Search google, bing, yahoo, and other search engines with python
Project description
search_engines
A Python library that queries Google, Bing, Yahoo and other search engines and collects the results from multiple search engine results pages.
Please note that web-scraping may be against the TOS of some search engines, and may result in a temporary ban.
Supported search engines
Google
Bing
Yahoo
Duckduckgo
Startpage
Aol
Dogpile
Ask
Mojeek
Brave
Torch
Features
- Creates output files (html, csv, json).
- Supports search filters (url, title, text).
- HTTP and SOCKS proxy support.
- Collects dark web links with Torch.
- Easy to add new search engines. You can add a new engine by creating a new class in
search_engines/engines/
and add it to thesearch_engines_dict
dictionary insearch_engines/engines/__init__.py
. The new class should subclassSearchEngine
, and override the following methods:_selectors
,_first_page
,_next_page
. - Python2 - Python3 compatible.
Requirements
Python 2.7 - 3.x with
Requests and
BeautifulSoup
Installation
Run the setup file: $ python setup.py install
.
Done!
Usage
As a library:
from search_engines import Google
engine = Google()
results = engine.search("my query")
links = results.links()
print(links)
As a CLI script:
$ python search_engines_cli.py -e google,bing -q "my query" -o json,print
Other versions
- async-search-scraper A really cool asynchronous implementation, written by @soxoj
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
No source distribution files available for this release.See tutorial on generating distribution archives.
Built Distribution
Close
Hashes for Search_Engines_Scraper_Tasos-0.5-py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 3b26c5cb318364edb2ca8640a2e7f8f9c98cb7ff58e7dcc674fa15f275d078d7 |
|
MD5 | 847a73ab2f7d68f14fe0780a7e6d720e |
|
BLAKE2b-256 | e5d316aff11bcedb90e5f14d49b894dce00c13a1ea8973e9aee59938cdf52f52 |