Skip to main content

DAVE, Data Acquisition and Validation Engine for production AI-powered scraping

Project description

DAVE

AI-powered scraping without LangChain, graph DSLs, or dependency bloat.

Give DAVE a URL and get clean, validated, structured data. DAVE ships with five runtime dependencies, no LangChain dependency, and no graph framework to learn. Use a prompt, a Pydantic schema, a built-in recipe, or nothing at all.

Tests PyPI Python versions Downloads License


Try it in 10 seconds

pip install dave-ai
dave extract "https://openai.com" --recipe company_info

Or use the library in one line:

import dave

result = await dave.extract("https://competitor.com")

That call has no prompt and no schema. DAVE still returns intelligent structured data by detecting the page type, important entities, key facts, calls to action, prices, contact signals, and links. This is the zero-config path for demos, exploration, and first-pass research.

The 3-line version

import dave

result = await dave.extract("https://example.com", "get the title and description")
print(result)

DAVE is designed for developers who need more than a prototype but do not want a full LangChain dependency tree. It combines fetcher selection, LLM extraction, Pydantic validation, caching, retries, rate limits, heuristic confidence scoring, cost tracking, batch jobs, and a plugin system in one typed Python package with only five runtime dependencies.

Why DAVE?

Most AI scraping tools are impressive in a notebook and heavy in a production job. ScrapeGraphAI publicly documents a graph-based API and a broad LangChain dependency stack. DAVE takes the opposite path: a small typed core, direct fetcher and extractor interfaces, and production controls that are easy to reason about.

What developers need How DAVE handles it
Keep dependency risk low Ships with five runtime dependencies: Beautiful Soup, HTTPX, Pydantic, Rich, and Typer
Avoid framework lock-in Uses no LangChain dependency and no graph DSL
Scrape static and JavaScript-heavy pages Auto-selects HTTP or Playwright, with extension points for Firecrawl and Crawl4AI
Extract with a prompt or a schema Accepts natural language prompts, Pydantic models, built-in recipes, or no prompt at all
Trust the output Validates schema extraction through Pydantic and reports transparent heuristic confidence
Keep costs visible Estimates tokens before runs and tracks cost after runs
Run hundreds of URLs Includes batch mode, retries, rate limits, progress, and JSON output
Avoid building glue code Ships caching, queueing, structured logging, CLI commands, and plugin registration

Confidence is a transparent heuristic based on evidence presence and source-text overlap, not a model-reported probability.

Feature comparison

This table is intentionally conservative. It compares publicly documented positioning and core APIs, not private roadmaps or unmeasured performance claims.

Capability DAVE ScrapeGraphAI Firecrawl Crawl4AI BeautifulSoup
Core positioning Lean Python extraction engine Graph-based AI scraping library Hosted web data API Open LLM-friendly crawler and scraper HTML and XML parser
LangChain dependency in core install No Yes, documented in package metadata Not applicable to hosted API usage No claim in docs reviewed No
Runtime dependency count in Python package Five core dependencies Broad dependency tree in package metadata SDK plus hosted service Framework package with browser-focused stack Parser library
Natural-language extraction Yes Yes, via graph APIs Yes, via Extract API Yes, via LLM extraction strategies No
Schema-shaped extraction Yes, Pydantic-first Yes, documented schema examples Yes, JSON schema in Extract API Yes, CSS, XPath, and LLM strategies Manual code
Zero-prompt, zero-schema URL extraction Yes, built into DAVE Not documented in sources reviewed Prompt or schema documented for Extract Not documented in sources reviewed Manual code
Built-in recipes for common pages Yes, company, pricing, jobs, contact, product, reviews Not a primary documented abstraction Not a primary documented abstraction Not a primary documented abstraction Manual code
Fetching model HTTP plus optional Playwright, plugin fetchers Playwright and provider integrations documented Hosted scrape, crawl, search, extract, interact API Browser crawler with advanced controls Bring your own HTTP client
Search plus extraction Yes, dave search over a pluggable provider SearchGraph documented Search and Extract documented Adaptive crawling documented Manual code
Best fit Lean structured extraction with production controls Graph-oriented AI scraping workflows Managed crawling and extraction API Configurable open crawler framework Deterministic HTML parsing

Sources reviewed: ScrapeGraphAI README, ScrapeGraphAI package metadata, Firecrawl docs, Firecrawl Extract docs, Crawl4AI docs, and Beautiful Soup docs.

Zero-config magic

DAVE can run with no prompt and no schema:

import dave

page = await dave.extract("https://stripe.com")

A typical result looks like this:

{
  "page_type": "company_or_content",
  "title": "Stripe",
  "summary": "Financial infrastructure for the internet",
  "key_entities": ["Stripe", "Payments", "Billing"],
  "key_facts": ["Stripe builds programmable financial services for businesses."],
  "links": ["https://stripe.com/pricing"],
  "contacts": {"emails": [], "phones": []},
  "prices": [],
  "products": ["Payments", "Billing", "Connect"],
  "jobs": [],
  "calls_to_action": ["Start now", "Contact sales"]
}

Zero-config extraction is not a replacement for schemas in critical workflows. It is the fastest way to inspect a new page, build lead lists, prototype competitive intelligence, and discover what schema you should use next.

Recipes that work out of the box

DAVE recipes are typed extraction flows for the web pages developers scrape every week.

Recipe Python CLI
Company info await dave.recipes.company_info(url) dave extract URL --recipe company_info
Pricing tiers await dave.recipes.pricing(url) dave extract URL --recipe pricing
Job listings await dave.recipes.job_listings(url) dave extract URL --recipe job_listings
Contact info await dave.recipes.contact_info(url) dave extract URL --recipe contact_info
Product features await dave.recipes.product_features(url) dave extract URL --recipe product_features
Reviews and testimonials await dave.recipes.reviews(url) dave extract URL --recipe reviews

Example:

import dave

company = await dave.recipes.company_info("https://linear.app")
pricing = await dave.recipes.pricing("https://linear.app/pricing")

The returned objects are Pydantic models, so they are easy to validate, serialize, and pass into downstream systems.

Beautiful CLI output

DAVE is built to look good in the terminal because the terminal is where scraping jobs fail, recover, and prove they are working.

$ dave extract "https://openai.com" --recipe company_info

╭──────────── Cost estimate ────────────╮
│ This extraction will use about        │
│ 2,000 tokens for roughly $0.003000.   │
╰───────────────────────────────────────╯
Proceed? [Y/n] y

╭──── Data Acquisition and Validation Engine ────╮
│ DAVE is extracting structured data              │
╰────────────────────────────────────────────────╯
+ Fetching page with auto fetcher 0:00:01
field name = OpenAI
field description = AI research and deployment company
field tech_stack = ["Python", "Kubernetes", "LLM APIs"]
field social_links = ["https://twitter.com/openai"]

╭────────────────── DAVE Extraction: company_info ──────────────────╮
│ Field        │ Value                                               │
│ name         │ OpenAI                                              │
│ description  │ AI research and deployment company                  │
│ founders     │ []                                                  │
│ funding      │ null                                                │
│ employees    │ null                                                │
│ tech_stack   │ ["Python", "Kubernetes", "LLM APIs"]              │
╰───────────────────────────────────────────────────────────────────╯

Run Metadata
confidence  0.91
cost_usd    0.003
fetcher     playwright
final_url   https://openai.com

Use machine-readable output when you need it:

dave extract "https://example.com" --prompt "get the title" --output json

Search the web, then extract

Skip the step where you go find URLs yourself. Give DAVE a query and it searches the web, then runs the normal extraction pipeline over each result.

dave search "best open source CRM" --recipe company_info --limit 5
import dave

report = await dave.search_extract("best open source CRM", prompt="get the company name and pricing", limit=5)
for item in report.ok_items:
    print(item.hit.url, item.data)

Search providers are pluggable so the core stays dependency-light. The default uses DuckDuckGo's HTML endpoint over the HTTPX and Beautiful Soup that DAVE already ships, so there is no API key and no extra dependency. Register your own provider for an internal search index or a paid search API:

from dave.plugins import register_search
from dave.search.base import BaseSearchProvider, SearchHit

class InternalSearch(BaseSearchProvider):
    name = "internal"

    async def search(self, query: str, *, limit: int = 5) -> list[SearchHit]:
        return [SearchHit(url="https://intranet.example/doc/1", title="Doc 1", rank=1)]

register_search("internal", InternalSearch())
dave search "quarterly targets" --search-provider internal --recipe company_info

Per-result failures are isolated: a URL that fails to fetch or extract is reported with its error while the rest of the batch continues.

Vision: structured data from images

When the DOM lies — canvas-rendered apps, image-only pages, scanned documents, screenshots, invoices — extract from the pixels instead. DAVE sends the image to a vision-capable model (OpenAI, Anthropic, or Gemini) and returns the same validated, confidence-scored structured data the text pipeline produces.

import dave
from pydantic import BaseModel

class Invoice(BaseModel):
    vendor: str
    total: str

invoice = await dave.extract_image("invoice.png", Invoice)
print(invoice.vendor, invoice.total)

Pass a file path or raw bytes. Image extraction has no source text to ground against, so confidence reflects field completeness and is labelled honestly rather than faked as source overlap.

Crawl a whole site

Point DAVE at a starting URL and it follows links across the site, extracting from every page. Breadth-first, bounded by page count and depth, same-domain by default. Every page goes through the full pipeline — cache, retries, rate limits, robots.txt, and confidence all apply per page.

dave crawl "https://example.com" --recipe company_info --max-pages 20 --max-depth 2
import dave

report = await dave.crawl("https://example.com", "get the page title and summary", max_pages=20, max_depth=2)
for page in report.ok_items:
    print(page.url, page.data)

A page that fails to fetch or extract is recorded with its error; the crawl keeps going. Use --any-domain to follow external links.

Batch mode

Process hundreds of known URLs with progress, retries, cache hits, and a cost summary.

dave batch urls.txt --recipe pricing --output results.json

The input file is plain text:

https://example.com/pricing
https://another-example.com/pricing
https://vendor.example/pricing

DAVE writes a JSON array with one item per URL. Failed URLs include the error message and successful URLs include data, confidence, evidence, cost, final URL, and fetcher metadata.

Cost calculator

DAVE estimates token usage before running an extraction.

This extraction will use about 2,000 tokens for roughly $0.003000.
Proceed? [Y/n]

The estimate uses the selected model, prompt, schema, and fetched page text. After extraction, DAVE returns tracked usage and cost metadata so teams can monitor spending by job, domain, recipe, or customer.

Streaming extraction

When --stream is enabled, DAVE emits field events as data is found instead of waiting until the full extraction is complete.

from dave.core.engine import DaveEngine

engine = DaveEngine()

async for event in engine.stream_extract("https://example.com", "get title and pricing"):
    print(event.type, event.message, event.data)

Streaming is useful for long pages, operator dashboards, live demos, and batch jobs where early partial output is better than silence.

Pydantic-first extraction

from pydantic import BaseModel, Field
import dave

class PricingTier(BaseModel):
    name: str
    price: str | None = None
    features: list[str] = Field(default_factory=list)

class PricingPage(BaseModel):
    tiers: list[PricingTier] = Field(default_factory=list)

pricing = await dave.extract("https://vendor.example/pricing", PricingPage)
print(pricing.model_dump())

Every schema extraction is validated before it is returned. If a provider returns invalid JSON or mismatched data, DAVE fails loudly instead of silently contaminating your pipeline.

Architecture

flowchart LR
    A[URL plus prompt, schema, recipe, or zero-config] --> B[DAVE Engine]
    B --> C{Fetcher Router}
    C --> D[HTTP Fetcher]
    C --> E[Playwright Fetcher]
    C --> F[Firecrawl or Crawl4AI Plugin]
    D --> G[Cache and Rate Limiter]
    E --> G
    F --> G
    G --> H[Chunker]
    H --> I{LLM Provider}
    I --> J[OpenAI]
    I --> K[Anthropic]
    I --> L[Ollama]
    H --> M[Pydantic Validator]
    J --> M
    K --> M
    L --> M
    M --> N[Confidence Scoring]
    N --> O[Structured Data]
    O --> P[CLI, Python API, Batch Jobs]

Fetchers

DAVE ships with a multi-fetcher backend and a simple routing model.

Fetcher Use it for Command
auto Let DAVE choose based on page signals --fetcher auto
http Fast static pages and APIs --fetcher http
playwright JavaScript-rendered pages --fetcher playwright
stealth Sites that block ordinary automation --fetcher stealth
file Local files: text, HTML, Markdown, JSON, PDF auto for local paths and file://
plugin Firecrawl, Crawl4AI, internal crawlers, or custom fetchers --fetcher your_name

Proxy rotation is configured through DaveConfig.proxies. Bring your own proxy URLs and DAVE will rotate them across requests.

Respecting robots.txt

Set respect_robots_txt=True (or DaveConfig(respect_robots_txt=True)) and DAVE checks each site's robots.txt before fetching, caches the result per domain, and refuses disallowed URLs with a clear RobotsDisallowedError. If robots.txt is unreachable, DAVE allows the fetch. Local files are never subject to robots rules.

Stealth fetcher

Some sites block headless browsers with bot walls and 403s. The stealth fetcher runs Playwright with an evasion layer — automation-control launch flags, a realistic browser context and client hints, and an init script that hides the signals most bot walls inspect (navigator.webdriver, missing window.chrome, headless WebGL vendor strings). It adds no heavy dependency such as undetected-playwright; it only needs Playwright, installed through the stealth extra.

pip install "dave-ai[stealth]"
python -m playwright install chromium
dave search "best open source CRM" --recipe company_info --fetcher stealth
import dave
from dave.core.config import DaveConfig
from dave.fetchers.stealth import register_stealth_fetcher

register_stealth_fetcher()
result = await dave.extract("https://example.com", "get the title", config=DaveConfig(fetcher="stealth"))

Stealth reduces detection; it is not a guarantee. Pair it with DaveConfig.proxies for sites that rate-limit or block by IP. Use it where you have permission to scrape, and respect each site's terms of service and robots.txt.

Plugin system

DAVE is designed for community extensions. Register custom fetchers and extractors without forking the project.

from dave.plugins import register_fetcher
from dave.fetchers.base import Fetcher, FetchResult

class InternalFetcher(Fetcher):
    name = "internal"

    async def fetch(self, url: str) -> FetchResult:
        return FetchResult(
            url=url,
            final_url=url,
            status_code=200,
            content="<html><title>Internal</title></html>",
            content_type="text/html",
            elapsed_ms=12.0,
            rendered=False,
        )

register_fetcher("internal", InternalFetcher)
dave extract "https://internal.example" --fetcher internal --prompt "summarize this page"

Plugins make it possible to add authenticated browser sessions, enterprise crawlers, proxy vendors, custom LLM routers, domain-specific validators, and private data enrichment steps.

Benchmarks coming

A reproducible benchmark suite is in development. We will publish real numbers, not targets, once it lands. Track it in issue #1.

The benchmark suite will measure end-to-end extraction from real product, pricing, jobs, documentation, and contact pages. It will publish fixtures, scripts, provider settings, hardware assumptions, network assumptions, and raw results so maintainers and users can reproduce the numbers.

Configuration

from dave.core.config import DaveConfig, LLMConfig

config = DaveConfig(
    fetcher="auto",
    llm=LLMConfig(provider="openai", model="gpt-4o-mini"),
    cache_path=".dave/cache.sqlite3",
    rate_limit_per_domain=2.0,
    proxies=["http://user:pass@proxy.example:8080"],
)

Environment variables are supported for common provider keys:

Variable Purpose
DAVE_LLM_PROVIDER Select openai, anthropic, gemini, groq, mistral, ollama, or mock
DAVE_LLM_API_KEY Provider API key (used by any provider)
OPENAI_API_KEY OpenAI fallback key
ANTHROPIC_API_KEY Anthropic fallback key
GEMINI_API_KEY / GOOGLE_API_KEY Gemini fallback key
GROQ_API_KEY Groq fallback key
MISTRAL_API_KEY Mistral fallback key
DAVE_LLM_MODEL Model override

All cloud providers run over plain HTTPX with native request shapes — no LangChain, no provider SDKs. OpenAI, Groq, and Mistral share the OpenAI-compatible chat API; Gemini and Anthropic use their native endpoints; Ollama runs locally.

Local development

git clone https://github.com/workwithlos-ui/dave.git
cd dave
python -m pip install -e ".[dev]"
pytest
ruff check .

Install Playwright support when you want JavaScript rendering:

python -m pip install -e ".[playwright]"
python -m playwright install chromium

The stealth fetcher uses the same Playwright runtime, available through the stealth extra:

python -m pip install -e ".[stealth]"
python -m playwright install chromium

Project structure

dave/
  core/          Engine, config, retries, queue, rate limits
  fetchers/      HTTP, Playwright, stealth, Firecrawl, Crawl4AI, router
  search/        Web search providers (DuckDuckGo, mock, pluggable)
  extractors/    LLM extraction, schema validation, confidence
  antibot/       Proxy and user-agent helpers
  cache/         SQLite response cache
  monitoring/    Cost tracking and structured logging
  cli/           Rich and Typer command line interface
  plugins.py     Community extension registry
  recipes.py     Built-in extraction recipes

Roadmap

Milestone Status
Zero-config extraction Done
Built-in recipes Done
Streaming CLI output Done
Batch mode Done
Plugin registry Done
Search plus extraction Done
Stealth fetcher plugin Done
Semantic chunking Done
Native multi-provider (OpenAI, Anthropic, Gemini, Groq, Mistral, Ollama) Done
Local file and PDF input Done
robots.txt politeness Done
Vision: structured data from images Done
Multi-page crawling Done
Hosted benchmark dashboard Planned
More recipe packs Planned
OpenTelemetry traces Planned

Contributing

DAVE is built for developers who scrape in production. Contributions should improve reliability, ergonomics, tests, documentation, or real-world extraction quality.

Before opening a pull request, please run:

pytest
ruff check .
python -m compileall dave

Good first contributions include new recipes, fetcher plugins, provider adapters, benchmark cases, docs examples, and domain-specific validators. Please keep code typed, keep tests meaningful, and do not add em dash characters to any file.

License

DAVE is released under the MIT License. See LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dave_ai-0.1.2.tar.gz (60.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dave_ai-0.1.2-py3-none-any.whl (61.1 kB view details)

Uploaded Python 3

File details

Details for the file dave_ai-0.1.2.tar.gz.

File metadata

  • Download URL: dave_ai-0.1.2.tar.gz
  • Upload date:
  • Size: 60.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.15

File hashes

Hashes for dave_ai-0.1.2.tar.gz
Algorithm Hash digest
SHA256 f5b8e3915060fc0b07012a0c261e6ef8d694744553e191d9e713500e7c7fe854
MD5 534998074e2869749178f7b41e57ce35
BLAKE2b-256 408a9b1b4d7a3239cb34bbd28e3f26d06817cfca0d5e0e2b7c3b1cae4a97e61a

See more details on using hashes here.

File details

Details for the file dave_ai-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: dave_ai-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 61.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.15

File hashes

Hashes for dave_ai-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 4edcdc6c6c9ff74c0d40f88a987d237924a1167490f3102ddbb7021a3909425f
MD5 c110a2036c2a9f990f715b927243a908
BLAKE2b-256 a05d7789c1d2f0ea2b0b6f1c5871addef7145c3f473753640d5812d1eb09bc29

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page