Skip to main content

OntoCast OntoCast logo

Agentic ontology-assisted extraction of RDF knowledge graphs from documents.

Python PyPI version PyPI Downloads Docs License pre-commit DOI

OntoCast turns unstructured text into queryable RDF: it co-evolves domain ontologies and fact graphs in a parallel map/reduce pipeline, with RDF 1.2 provenance, entity disambiguation across chunks, and optional vector-backed ontology retrieval. Run it as a REST service, a batch CLI, or embed the pipeline in your own LangChain / LangGraph agent.

Documentation: growgraph.github.io/ontocast


Why OntoCast

Most extractors dump triples and leave ontology drift to you. OntoCast treats schema and instance data as one loop: per-chunk render → critic → merge, with GraphUpdate patches (insert/delete) instead of regenerating whole graphs, SHACL validation with LLM-free autofix, and a light install so you can embed the core without pulling Docling, gRPC, or ONNX.


Features

  • Parallel ontology + facts loops — concurrent per-unit render/critic with configurable workers
  • GraphUpdate patches — token-efficient insert/delete ops, not full-graph regeneration
  • Entity disambiguation — embedding + symbolic alignment across chunks
  • RDF 1.2 provenance — quoted triples / provenance artifacts; optional strip_provenance
  • Ontology context — catalog selection, vector retrieval (LanceDB or Qdrant), or a fixed ontology
  • Facts validation — invariants, SHACL, and machine repairs without an extra LLM pass
  • Stores — in-memory pyoxigraph by default; Fuseki for persistence; tenancy by tenant/project
  • LLM caching — disk cache, in-flight limits, optional read-only / batch pre-warm
  • Embeddable — ontocast_tools, run_unit_pipeline, or a LangGraph node

Install

Pick at least one LLM provider extra. Add server for the CLI and HTTP API:

uv add "ontocast[server,openai]"
# or: pip install "ontocast[server,openai]"

Common add-ons: doc-processing (PDF/DOCX), lancedb or qdrant (ontology retrieval), shacl (shape validation).

uv add "ontocast[server,openai,doc-processing,lancedb,shacl]"

Full extras table: Installation.


Quick start

cp .env.example .env
# Set LLM_API_KEY (and LLM_PROVIDER / LLM_MODEL_NAME as needed)

ontocast serve
curl -X POST http://localhost:8999/process -F "file=@document.pdf"

Batch without a server:

ontocast process --input-path ./document.pdf --head-chunks 5 --output-dir ./out

Omit FUSEKI_URI for in-memory pyoxigraph. Details: Quick Start.

Supplying Your Ontologies

OntoCast uses seed ontologies (in Turtle .ttl format) to guide extraction. Provide yours in two ways:

  1. Directory Seed: Set ONTOCAST_ONTOLOGY_DIRECTORY=/path/to/your/ontologies in your .env. All .ttl files in that folder sync automatically on startup.
  2. API Upload: Register schemas dynamically with the running server:
    curl -X POST "http://localhost:8999/ontologies?tenant=ontocast&project=test" -F "file=@my_ontology.ttl"
    

Configuration

Start from .env.example.minimal — 47 variables instead of 202, grouped by the decision they belong to. Then pick a playbook for what you are actually doing: evaluating, building an ontology, populating facts, scaling to a large catalog, or serving it.

The knobs that change what the pipeline does — as opposed to where it stores things:

Variable Default What it controls
RENDER_MODE ontology_and_facts Which halves run. ontology writes no facts; facts skips the ontology block and extracts only against the catalog you already have
ONTOLOGY_CONTEXT_MODE selected_single_ontology Where each unit's schema comes from: LLM catalog selection, vector retrieval, or one pinned ontology
LLM_GRAPH_FORMAT jsonld Wire encoding the LLM emits graphs in; turtle is the legacy alternative
MAX_VISITS_PER_NODE 1 Render/critic retry budget. At 1 the LLM critic never runs
PARALLEL_WORKERS 16 Concurrent content-unit workers
LLM_PROVIDER / LLM_MODEL_NAME / LLM_API_KEY openai Provider selection and credentials
ONTOCAST_ONTOLOGY_DIRECTORY — Seed ontologies synced on startup
FUSEKI_URI — Triple store; unset means in-memory pyoxigraph

RENDER_MODE, ONTOLOGY_CONTEXT_MODE and LLM_GRAPH_FORMAT are also per-request parameters on /process. Full surface, including chunking, retrieval and validation: Configuration.


Embed in your agent

from langchain.agents import create_agent
from ontocast import Config, ToolBox, ontocast_tools

tools = await ToolBox.acreate(Config.in_memory())
await tools.initialize()

agent = create_agent(
    model,
    tools=[*ontocast_tools(tools)],
    prompt="Edit the ontology from the user's text.",
)

Also: run_unit_pipeline for a single passage, or make_ontocast_node inside your own LangGraph — see Embedding OntoCast.


Workflow

Workflow diagram

  1. Convert → chunk prepare (segment, tag, filter, size)
  2. Parallel ontology render → normalize → consolidate → structural check → critic
  3. Parallel facts render → merge / disambiguate → validate (invariants, SHACL, autofix)
  4. Serialize to the triple store; return Turtle from the API

Workflow guide · landscape: graph.lr.png · per-unit: ontology_loop, facts_loop


Documentation

Everything lives at growgraph.github.io/ontocast:

Installation · Quick Start Getting started
Core Concepts · Workflow · Configuration How it works
API · Embedding · Tenancy Integrate
Ontology Context · Validation / SHACL · Triple Stores Operate
API Reference Python API

Release notes: CHANGELOG.md


Contributing

See Contributing. Issues and discussion: GitHub.

License

Apache License 2.0 — see LICENSE.

Metadata

Release files for ontocast 0.6.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ontocast 0.6.2
File Size Uploaded
ontocast-0.6.2.tar.gz 926.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ontocast 0.6.2
File Interpreter ABI Platform
ontocast-0.6.2-py3-none-any.whl Python 3 none any Details

Total release size: 1.6 MB

Release files / ontocast-0.6.2.tar.gz

Download URL ontocast-0.6.2.tar.gz
Size 926.6 kB
Tags Source
SHA-256 checksum
How to use checksums
c005c24a8fb1afb0c4f8119a8b81257e2c899eeb7fca0a855106e36aabdc376a
BLAKE2b-256 checksum
How to use checksums
e61e9117b87a8882bc72a1fb7a93eebdf7006496fb0d3f851fe3d93f2688f6de
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.12 {"installer":{"name":"uv","version":"0.10.12","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / ontocast-0.6.2-py3-none-any.whl

Download URL ontocast-0.6.2-py3-none-any.whl
Size 651.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cfe465d9ac0327577973de87f5fcc9f7bb9dacde283810232821d71472f5b41f
BLAKE2b-256 checksum
How to use checksums
39a17f7ee8e5109b4a25363fd6efbb1acc5249ec9d41ea0c4100e8d0686c0944
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.12 {"installer":{"name":"uv","version":"0.10.12","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.6.7

2 release files

0.6.6

2 release files

0.6.5

2 release files

0.6.4

2 release files

This release

0.6.2 This release

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.4.3

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page