langchain-opensolr
LangChain integration for Opensolr — managed Apache Solr with server-side embeddings and native hybrid (BM25 + kNN) search.
Product page: opensolr.com/langchain · Platform: opensolr.com (managed Solr hosting since 2011 — free 15-day trial, no card)
No local embedding model. No third-party embedding API key. One set of credentials, and the vectors are computed on Opensolr's GPU infrastructure (multilingual E5-large-instruct, 1024 dimensions, cosine).
pip install langchain-opensolr
The whole tutorial
from langchain_opensolr import OpensolrVectorStore
vs = OpensolrVectorStore(
index="mysite__dense", # vector-enabled Opensolr index
email="you@example.com",
api_key="YOUR_OPENSOLR_API_KEY",
create_if_missing=True, # provisions the index on first use
)
vs.add_texts(
["Hybrid search fuses BM25 with vector similarity",
"Cats sleep sixteen hours a day"],
metadatas=[{"category": "search"}, {"category": "animals"}],
)
docs = vs.similarity_search("how do lexical and semantic search combine?", k=1)
print(docs[0].page_content)
That's it — no embedding model was configured, because embedding happens on the server at both index and query time.
Hybrid search
Pure vector search fails on exact identifiers; pure BM25 fails on meaning.
Opensolr's {!hybrid} query parser fuses both scores per document:
docs = vs.similarity_search(
"affordable restaurants",
k=5,
hybrid=True,
mode="union", # union | keywords_required | meaning_required | intersection
alpha=0.5, # 0 = all semantic … 1 = all lexical
)
Metadata filters
vs.similarity_search("search engines", k=5, filter={"category": "search"})
vs.similarity_search("anything", k=5, filter='meta_rank:[2 TO *]') # raw Solr fq
Metadata round-trips losslessly (stored as JSON alongside filterable
meta_* fields).
As a retriever, in any chain
retriever = vs.as_retriever(search_kwargs={"k": 5, "hybrid": True})
Standalone embeddings
Use Opensolr's embedding endpoint with any other LangChain component:
from langchain_opensolr import OpensolrEmbeddings
emb = OpensolrEmbeddings(email="you@example.com", api_key="...", index="mysite__dense")
emb.embed_query("budget-friendly dining") # -> 1024 floats
Notes
- Vector-enabled indexes run on Opensolr's Solr 9.x environments — currently
us(Chicago),de(Germany),fi(Finland). Passlocation=to choose. The list is fetched live from the platform, so new regions work without a package upgrade — and additional dedicated regions can be deployed on request (paid add-on): support@opensolr.com. - A free Opensolr account (15-day trial, no card) includes an AI quota that comfortably covers this README end to end: opensolr.com.
- Full platform docs: AI & Vector Search.
Development
pip install -e . pytest
pytest tests/unit_tests
OPENSOLR_EMAIL=... OPENSOLR_API_KEY=... OPENSOLR_INDEX=... pytest tests/integration_tests
How writing works (Data Ingestion API)
Writes go through Opensolr's Data Ingestion API
— the same pipeline the Drupal and WordPress connectors use. It is
asynchronous: documents are queued, then embeddings, sentiment, language
and all crawler-identical derived fields are computed server-side, and
documents become searchable within about a minute. Progress is visible in the
Opensolr Control Panel and via the ingest_status API. Each document's
identity is its uri (the Solr id is md5(uri)): pass a real URL in
metadata ({"uri": "https://..."}), or a deterministic one is synthesized
from your id. Re-submitting the same uri updates the document. Pass
{"rtf": True, "uri": "https://.../file.pdf"} and the server extracts the
text from PDF/DOCX/XLSX for you.
Lexical-only mode
Don't need vectors? Pure keyword search skips the embedding call entirely — zero AI quota, and it works on any Opensolr index, including non-vector ones and older Solr versions.
Your index schema
Documents follow the Opensolr document model (title, description, text,
meta_* custom fields). To see the full schema: Control Panel → click your
index → Configuration → Edit File → schema.xml. Prefer zero-effort data
entry? Configure the Web Crawler in the Control Panel (Index Tools →
WebCrawler): add your site URL, validate it, and Opensolr indexes the whole
site for you.
How it's tested
Every release is validated against live Opensolr infrastructure — no mocks:
- Unit tests (offline): location aliases, filter→fq mapping, query building, escaping.
- End-to-end suite: the full write path through the async Data Ingestion
queue (queued → server-side enrichment → searchable), semantic / hybrid /
lexical retrieval, metadata round-trip, filters, id round-trip (your ids
and the Solr
md5(uri)ids), deletes by id and by query. - Real-corpus validation: searches run against a 340-document replica of opensolr.com's own production search index. Verified: pure-semantic hits with zero keyword overlap ("how do I get my data back after a disaster" → backup & restore docs), cross-lingual queries (Romanian query → English content), exact-term surfacing in hybrid mode, all four hybrid modes, and the full alpha range 0 → 1.
- PDF ingestion: a real PDF ingested via
rtf:true— server-side text extraction (13k+ chars), automatic content-type detection, then retrieved with a purely semantic query against its contents.
pytest tests/unit_tests
OPENSOLR_EMAIL=... OPENSOLR_API_KEY=... OPENSOLR_INDEX=... pytest tests/integration_tests
MIT license.
Release files for langchain-opensolr 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| langchain_opensolr-0.2.1.tar.gz | 16.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| langchain_opensolr-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 32.9 kB
Release files / langchain_opensolr-0.2.1.tar.gz
| Download URL | langchain_opensolr-0.2.1.tar.gz |
|---|---|
| Size | 16.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
aef2b77100525e714c4c3976e108a00e3c29568c8caa258ec4231ba85689fa22
|
|
BLAKE2b-256 checksum How to use checksums |
2f54dcfa76311561b28274e627e629319681d467c137b5804f950d59ee4c3709
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.12
|
Release files / langchain_opensolr-0.2.1-py3-none-any.whl
| Download URL | langchain_opensolr-0.2.1-py3-none-any.whl |
|---|---|
| Size | 16.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
83eb8b0aa0f9252ccb7f356a03f2b54aaa02de9a25063a7b42c5972229964aa6
|
|
BLAKE2b-256 checksum How to use checksums |
d96352acdc60f6c95ff953f6c50785441e56280a6653a5d69852fb4abd6ddbd9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.12
|