🕷️ Scrapy → Meilisearch Pipeline
A Scrapy pipeline that batches items and indexes them into Meilisearch, with optional index creation and index settings using the modern Meilisearch Python client.
✨ Features
- ✅ Uses the official modern Meilisearch client (TaskInfo / Pydantic models)
- 🧰 Optional index creation (with
primaryKey) and index settings update - 📦 Batching of items before insertion
- 🔎 Task tracking with status check (failed tasks are logged and stored)
- 🧪 Example Scrapy project + Docker Compose for Meilisearch
- 🧹 Tooling:
uv,pytest,ruff,black,mypy,justtasks
🧠 How batching works (pipeline logic)
The pipeline keeps two internal buffers:
_buffer→ a list of items waiting to be sent to Meilisearch_tasks→ a list of Meilisearch TaskInfo objects created byadd_documents()andupdate_settings()
Flow:
process_itemconverts an item todictand pushes it into_buffer.- When
_bufferlength reachesMEILI_BATCH_SIZE, the pipeline performs a flush:- Sends the whole
_bufferwithindex.add_documents(batch) - Appends the returned TaskInfo to
_tasks - Calls
_check_all_tasks(): waits on all tasks in_tasksviawait_for_task()and- if any task ends with
status="failed", it is moved to_failed_tasks - otherwise it is discarded (success) —
_tasksis cleared
- if any task ends with
- Sends the whole
close_spider:- If
_bufferstill has items, a final flush is executed (and tasks checked) - If
_tasksstill contains tasks (e.g., settings only), they are checked - If any failed tasks were detected, they are logged (no exception is raised by design)
- If
Benefits of this approach:
- Minimal memory use (bounded by
MEILI_BATCH_SIZE) - Early surfacing of Meilisearch task failures during the crawl
- Predictable and simple control flow
📦 Installation
From PyPI:
pip install scrapy-meili-pipeline
Using uv:
uv add scrapy-meili-pipeline
⚙️ Settings
Add the pipeline to Scrapy and configure Meilisearch via settings:
ITEM_PIPELINES = {
"scrapy_meili_pipeline.MeiliSearchPipeline": 300,
}
MEILI_URL = "http://127.0.0.1:7700"
MEILI_API_KEY = "masterKey" # or None
MEILI_INDEX = "articles" # required
MEILI_PRIMARY_KEY = "id" # optional
MEILI_INDEX_SETTINGS = { # optional
"filterableAttributes": ["author", "categories", "keywords", "rating"],
"sortableAttributes": ["published_at", "rating"],
"searchableAttributes": ["title", "summary", "content", "keywords"],
}
MEILI_BATCH_SIZE = 500
MEILI_TASK_TIMEOUT = 180
MEILI_TASK_INTERVAL = 1
This library supports ONLY the modern Meilisearch client and expects TaskInfo objects with a
task_uidattribute.
🚀 Quick example (Scrapy spider)
class ArticleSpider(Spider):
name = "articles"
custom_settings = {
"MEILI_INDEX": "news",
"MEILI_BATCH_SIZE": 200,
"MEILI_INDEX_SETTINGS": {"filterableAttributes": ["site", "tags"]},
}
def parse(self, response):
yield {
"id": response.url,
"title": response.css("h1::text").get(),
"author": response.css(".author::text").get(),
"content": response.css("article::text").getall(),
"rating": 4,
}
🧪 Example project & Meilisearch (examples/)
This repo ships with a runnable example under examples/ that scrapes the public test site
https://webscraper.io/test-sites/e-commerce/allinone and indexes product tiles into Meilisearch.
Start Meilisearch with Docker
cd examples
docker compose up -d
Meilisearch UI: http://127.0.0.1:7700
Run the example spider (via Just)
From the repository root:
just example
What the example task does:
- switches to
examples/simple_project - runs
scrapy crawl demo -s LOG_LEVEL=INFO
If you prefer running it manually:
cd examples/simple_project
uv run scrapy crawl demo -s LOG_LEVEL=INFO
🧱 Project structure
scrapy-meili-pipeline/
├── src/
│ └── scrapy_meili_pipeline/
│ ├── __init__.py
│ └── meili_pipeline.py
├── tests/
│ └── test_pipeline.py
├── examples/
│ ├── README.md
│ ├── .env.example
│ ├── docker-compose.meilisearch.yml
│ └── simple_project/
│ ├── scrapy.cfg
│ └── simple_project/
│ ├── __init__.py
│ ├── settings.py
│ ├── sitecustomize.py
│ └── spiders/
│ └── demo_spider.py
├── Justfile
├── pyproject.toml
├── README.md
├── LICENSE
└── .github/
└── workflows/
├── ci.yml
└── publish.yml
🛠️ Development
Using uv + just:
just sync # install all deps (dev included)
just check # ruff + black --check + mypy + pytest
just test # run unit tests
just coverage # terminal coverage
just coverage-html # HTML coverage at ./htmlcov/index.html
just build # build wheel + sdist (uv build)
just publish # publish to PyPI (uv publish)
Manual (without just):
uv sync --all-extras --dev
uv run ruff check .
uv run black --check .
uv run mypy .
uv run pytest
uv run pytest --cov=src --cov-report=html
uv build
uv publish
📜 License
Released under the MIT License.
Metadata
Release files for scrapy-meili-pipeline 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| scrapy_meili_pipeline-0.1.1.tar.gz | 11.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| scrapy_meili_pipeline-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 19.8 kB
Release files / scrapy_meili_pipeline-0.1.1.tar.gz
| Download URL | scrapy_meili_pipeline-0.1.1.tar.gz |
|---|---|
| Size | 11.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
13a49f98a6e297f628517759c5bf978658fc2838afe8cb11b946b4bdb63a3d8d
|
|
BLAKE2b-256 checksum How to use checksums |
a7733a7e3a0a690d94171ee134f26e564bf872c46eab285dce00c2afa58b9f8a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.9.5
|
Release files / scrapy_meili_pipeline-0.1.1-py3-none-any.whl
| Download URL | scrapy_meili_pipeline-0.1.1-py3-none-any.whl |
|---|---|
| Size | 8.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
aaa2643befd75a1430f28a5259a1ab635391dda71950b5c2ebc31e1a78a16236
|
|
BLAKE2b-256 checksum How to use checksums |
066d669d20d79f428dfc23364a88efa98deba2a4db5870502badbf253b43e004
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.9.5
|