BIE — BitSearch Intelligence Engine
A real-time web search and crawling toolkit for AI applications — no API keys, no subscriptions, no third-party search services.
BIE gives any LLM, RAG pipeline, or AI agent five core primitives — search, extract, map, crawl, and a hybrid index — all running locally on top of **BitS **, our async crawling framework. Use it as a Python library, REST API, CLI, or MCP server.
import bie
# Search the live internet — no URLs, no API key, no subscription
results = bie.websearch("latest semiconductor export rules 2026")
for r in results:
print(r.title, "—", r.url, f"(score={r.score:.3f})")
print(r.snippet)
# Get clean markdown from a specific page
page = bie.extract("https://example.com/article")
print(page.markdown)
Honest scope
BIE is built to be a genuinely useful, self-hosted web search/extraction toolkit — and we'd rather be upfront about what that means than oversell it:
- What's real: working search (free public discovery + Bitscrape crawl + hybrid BM25/vector ranking with query fan-out), Markdown extraction with JS-rendering fallback, sitemap-based site mapping, instruction-guided crawling, a prompt-injection heuristic scanner, and REST/CLI/MCP/LangChain integrations — all of it runs today, with no paid dependencies.
- What it isn't: a replacement for web-scale search infrastructure. BIE doesn't have its own crawled index of the internet — discovery relies on free public search endpoints (which can rate-limit), and relevance ranking is BM25+embeddings, not a model tuned on years of query logs. "Crawl guided by natural language" means keyword-relevance link prioritization, not an LLM reading every page. The prompt-injection scanner is a pattern-matching heuristic, not a guarantee.
If your use case needs guaranteed uptime, massive scale, or state-of-the-art ranking, a commercial search API may still be the right choice for that piece. BIE is for teams that want a capable, free, self-hosted starting point — and full control over the code.
Core primitives
| Function | What it does |
|---|---|
bie.websearch(query) |
Search the live internet — no URLs needed. Free discovery (DuckDuckGo + Bing fallback) with query fan-out, crawled and ranked by BIE's hybrid index. |
bie.extract(url) |
Fetch a URL and return clean Markdown, with nav/ads/scripts stripped. Optional JS rendering via Playwright. |
bie.map_site(url) |
Discover a site's sitemap(s) and the URLs they list, before crawling. |
bie.crawl_site(urls, instruction=...) |
Crawl a site, prioritizing links by keyword-relevance to your instruction. Returns an index + ranked results. |
bie.search(query, urls=...) |
Crawl specific URLs and rank their content against a query. |
bie.BIE() |
Build a persistent, queryable hybrid index across multiple crawls. |
bie.scan_for_prompt_injection(text) |
Heuristic scan for prompt-injection patterns in crawled content. |
Install
pip install bits-bie
Note: the PyPI distribution is named
bits-bie(sincebiewas too similar to an existing PyPI project), but you stillimport bieand run thebieCLI command — same API as shown below.
Optional extras:
pip install "bits-bie[embeddings]" # semantic/vector search (sentence-transformers)
pip install "bits-bie[server]" # FastAPI + Uvicorn REST server
pip install "bits-bie[mcp]" # Model Context Protocol server
pip install "bits-bie[render]" # JS rendering for extract() via Playwright
pip install "bits-bie[langchain]" # LangChain tool adapters
pip install "bits-bie[notebook]" # smoother async behaviour in Jupyter/Colab
pip install "bits-bie[all]" # everything
BIE depends on
bitscrape, our proprietary async crawling & extraction framework, which is installed automatically.
Usage
1. Search the live internet — no URLs, no API key, no subscription
import bie
results = bie.websearch("who won the latest F1 race")
for r in results:
print(r.title, "—", r.url)
print(r.snippet)
websearch pipeline:
- Discovery — free, public, no-key search endpoints (DuckDuckGo,
with an automatic Bing fallback). By default, several phrasings of
your query are searched and merged (
fanout=True) for better recall. - Crawl — discovered URLs are crawled with Bitscrape.
- Rank — extracted content is chunked and ranked against your query with BIE's hybrid BM25 + vector index.
- Security filter — results whose matched text trips the
prompt-injection heuristic (
bie.security) are dropped by default.
Useful options: top_k, discovery_results, fanout,
max_query_variants, deep, scan_security, use_embeddings.
2. Extract — clean Markdown from a specific URL
page = bie.extract("https://example.com/article")
print(page.title)
print(page.markdown)
print(page.word_count)
# For JS-rendered (SPA) pages:
page = bie.extract("https://app.example.com", render_js=True) # requires bie[render]
If a static fetch returns suspiciously little text, extract raises
ExtractError suggesting render_js=True rather than silently returning
near-empty content.
Every result includes page.security — a SecurityReport flagging
prompt-injection-like patterns in the extracted text (see
Security below).
3. Map — discover a site's structure before crawling
sitemap = bie.map_site("https://example.com")
print(sitemap.sitemap_urls) # which sitemap files were found
print(len(sitemap.urls)) # how many pages they list
print(sitemap.filter(r"/blog/")) # just the blog URLs
Based on the sitemaps.org protocol: reads robots.txt for Sitemap:
directives, falls back to /sitemap.xml, and recursively expands sitemap
indexes.
4. Crawl — guided by a natural-language instruction
engine, results = bie.crawl_site(
["https://docs.example.com"],
instruction="authentication and rate limits",
max_pages=30,
max_depth=2,
)
for r in results:
print(r.title, r.url)
# Re-query the same crawled index without re-crawling:
more = engine.search("error codes")
Outgoing links are ranked by keyword overlap between your instruction and each link's anchor text + URL path — a fast heuristic that biases the crawl toward relevant pages without an LLM call per page.
5. Search specific sites (no live-web discovery)
results = bie.search("AI regulation news", urls=["https://example.com/news"], top_k=5)
for r in results:
print(r)
6. Build a reusable index
from bie import BIE
engine = BIE()
engine.crawl(["https://example.com/blog", "https://another-site.com"])
print(engine.search("quarterly earnings"))
print(engine.search("product launch")) # reuses the same index
# Index your own text (no crawling):
engine.add_text(url="internal://doc-1", title="Q2 Memo", text="...", trust_score=1.0)
7. CLI
# Search the live internet — no URLs needed
bie search-live "who won the latest F1 race"
# Clean markdown from a URL
bie extract https://example.com/article
# Discover a site's sitemap
bie map https://example.com --filter "/blog/"
# Crawl, guided by an instruction
bie crawl https://docs.example.com --instruction "authentication and rate limits" --max-pages 30
# Crawl + search specific sites in one command
bie search "global markets today" --url https://www.bbc.com/news --top-k 5
# Run the REST API
bie serve --port 8000
# Run as an MCP server (stdio)
bie mcp
8. REST API
bie serve --port 8000
curl -X POST http://localhost:8000/search/live \
-H "Content-Type: application/json" \
-d '{"query": "who won the latest F1 race", "top_k": 5}'
curl -X POST http://localhost:8000/extract \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/article"}'
curl -X POST http://localhost:8000/map \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com"}'
curl -X POST http://localhost:8000/crawl/url \
-H "Content-Type: application/json" \
-d '{"urls": ["https://example.com/news"], "instruction": "pricing pages"}'
curl -X POST http://localhost:8000/search \
-H "Content-Type: application/json" \
-d '{"query": "latest news", "top_k": 5}'
See the full endpoint contract in docs/API.md.
9. MCP (Model Context Protocol)
Add BIE as a tool in your MCP client (e.g. claude_desktop_config.json):
{
"mcpServers": {
"bie": {
"command": "bie",
"args": ["mcp"]
}
}
}
This exposes six tools to your AI assistant:
bie_web_search(query, top_k, deep)— search the live internet, no URLs neededbie_extract(url, render_js)— fetch a URL as clean Markdownbie_map(url, filter_pattern)— discover a site's sitemapbie_search(query, urls, top_k, max_pages)— crawl + search specific URLsbie_crawl(urls, max_pages, instruction)— crawl & index into a session-persistent storebie_index_search(query, top_k)— search the session index
10. LangChain
from bie.integrations.langchain import get_tools
tools = get_tools() # [bie_websearch, bie_extract, bie_crawl_site]
# pass `tools` to your LangChain/LangGraph agent
Requires pip install "bits-bie[langchain]".
Security
BIE includes bie.scan_for_prompt_injection(text) — a pattern-based
heuristic that flags text likely to contain instructions aimed at an LLM
(e.g. "ignore previous instructions...", fake SYSTEM: blocks, requests
to reveal a system prompt).
bie.extract()attaches aSecurityReportto every result (result.security).bie.websearch()drops results whose matched chunk trips the heuristic by default (scan_security=True).
This is a signal, not a guarantee. It catches common, unobfuscated
injection phrasing in crawled web content — it will not catch everything,
and legitimate pages discussing prompt injection may occasionally be
flagged. Treat flagged=True as "review before feeding this directly
into a high-privilege agent context," not as "this content is dangerous"
or "unflagged content is safe." See bie/security.py for the full pattern
list and caveats.
Configuration
All settings can be set via environment variables prefixed with BIE_,
or passed directly:
from bie import BIE, BIESettings
engine = BIE(BIESettings(
max_pages=20,
max_depth=1,
use_embeddings=True,
embedding_model="sentence-transformers/all-MiniLM-L6-v2",
bm25_weight=0.6,
vector_weight=0.4,
))
| Setting | Env var | Default | Description |
|---|---|---|---|
max_pages |
BIE_MAX_PAGES |
40 |
Max pages crawled per seed URL |
max_depth |
BIE_MAX_DEPTH |
2 |
Max link-follow depth |
concurrent_requests |
BIE_CONCURRENT_REQUESTS |
16 |
Crawl concurrency |
robotstxt_obey |
BIE_ROBOTSTXT_OBEY |
true |
Respect robots.txt |
use_embeddings |
BIE_USE_EMBEDDINGS |
true |
Enable semantic search |
chunk_size |
BIE_CHUNK_SIZE |
800 |
Chars per chunk |
bm25_weight / vector_weight |
BIE_BM25_WEIGHT / BIE_VECTOR_WEIGHT |
0.5 / 0.5 |
Fusion weights |
discovery_backends |
BIE_DISCOVERY_BACKENDS |
ddg_html,ddg_lite,bing_html |
Ordered, comma-separated discovery backends for websearch(). Add searxng for a self-hosted instance. |
searxng_url |
BIE_SEARXNG_URL |
None |
Base URL of a self-hosted SearXNG instance, used by the searxng discovery backend |
api_key |
BIE_API_KEY |
None |
If set, requires Authorization: Bearer <key> |
Troubleshooting
TypeError: '<' not supported between instances of 'Request' and 'Request'
during a crawl — this was a Bitscrape scheduler bug (its priority queue
compared Request objects directly when two requests shared the same
priority). BIE patches bitscrape.Request to be orderable at import
time, so this no longer occurs. If you still see it, you're likely on an
older bits-bie version — upgrade.
RuntimeError: asyncio.run() cannot be called from a running event loop — Jupyter/Colab/IPython already run an event loop, which used to
break engine.crawl(urls) / bie.websearch(...). Both now detect a
running loop automatically and either use
nest_asyncio (install via
pip install "bits-bie[notebook]") or fall back to running the crawl on
a background thread — no code changes needed. If you're already inside
an async def, you can also call await engine.acrawl(urls) directly.
bie.websearch(...) returns [] / all discovery backends fail —
discovery scrapes DuckDuckGo/Bing's public HTML result pages, which can
be blocked or rate-limited. Call
bie.discovery.get_last_discovery_diagnostics() right after to see why:
import bie
from bie.discovery import get_last_discovery_diagnostics
results = bie.websearch("...")
if not results:
print(get_last_discovery_diagnostics().summary())
This distinguishes three cases:
- Network blocked — every backend failed at the connection level
(or an egress proxy returned
x-deny-reason: host_not_allowed). This environment can't reach these hosts at all — check its outbound network/proxy/firewall config. Common in sandboxed code-execution environments; Colab and most servers have unrestricted outbound access. - Blocked / rate-limited — backends responded with
403/429/etc., typically from bot-detection on a shared IP. Retry later, reduce request volume, or configure asearxngbackend (below). - Empty response — got
200 OKbut no parseable results (often a CAPTCHA/consent page).
For the most reliable no-API-key discovery, self-host SearXNG and add it as a backend:
export BIE_DISCOVERY_BACKENDS=searxng,ddg_html,ddg_lite,bing_html
export BIE_SEARXNG_URL=http://localhost:8080
Architecture
┌─────────────────────────────────────────────────┐
│ bie │
│ │
query ──▶ │ discovery (DuckDuckGo/Bing) ──▶ query fan-out │
│ │ │
urls ──▶ │ ▼ │
│ Crawler (Bitscrape) ──▶ Document ──▶ Chunker │
│ │ │ │
│ │ ▼ │
│ │ HybridIndex │
│ │ BM25 + Vector │
│ │ (RRF fusion) │
│ ▼ │ │
│ extract()/map() Ranked SearchResults │
│ (standalone) │ │
│ security scan │
└─────────────────────────────────────────────────┘
│ │ │ │
Python API REST API MCP Server LangChain
This OSS edition implements the core of the BIE PRD's Module 1
(Crawler), Module 2 (Indexes), Module 3 (Hybrid Retriever), and
Module 11 (Agent API) as a single lightweight package — no external
services required. Larger deployments can swap BM25Index/VectorIndex
for Elasticsearch/Milvus-backed implementations behind the same
HybridIndex interface.
Built on Bitscrape
BIE's crawling and extraction layer is powered by
BitS
(pip install bitscrape), our async, robots.txt-aware web scraping
framework — giving BIE high-performance, polite crawling out of the box.
License
MIT — see LICENSE.
Release files for bits-bie 1.2.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| bits_bie-1.2.4.tar.gz | 61.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| bits_bie-1.2.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 120.1 kB
Release files / bits_bie-1.2.4.tar.gz
| Download URL | bits_bie-1.2.4.tar.gz |
|---|---|
| Size | 61.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3e45f0e2d504984b9c02233c79657fab979ccf851bde354156d64b2411feace5
|
|
BLAKE2b-256 checksum How to use checksums |
8915139af138f4a0b914bf64a3bdaf600757096caf0d80613c3b7af39df85efe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 15, 2026.
Transparency logRelease files / bits_bie-1.2.4-py3-none-any.whl
| Download URL | bits_bie-1.2.4-py3-none-any.whl |
|---|---|
| Size | 58.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5a64c25f5fc7485a4b742976792a89f5ef7b0ef64656a9dc8c73b8e3bdf99c44
|
|
BLAKE2b-256 checksum How to use checksums |
23f5e716c341f99ef50c5e2ea960e67db568f44efae3c6128fc95bbb0959ae03
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 15, 2026.
Transparency log