Skip to main content

OntoCast OntoCast logo

Ontology-guided extraction of RDF knowledge graphs from documents.

Python PyPI version PyPI Downloads Docs License pre-commit DOI

OntoCast reads documents and writes an RDF knowledge graph: an ontology that describes the domain, and the facts the documents state in its terms. Give it your ontologies and it extracts facts against them; give it none and it builds one as it reads. Run it as an HTTP service, as a batch command, or inside your own LangChain or LangGraph agent.

Documentation: growgraph.github.io/ontocast

Features

  • Ontology and facts together. Each part of a document goes through a language model in a render-and-critique loop, in parallel; ontology changes are merged, versioned and checked before facts are extracted against them.
  • Patches, not rewrites. The model emits insert/delete updates to a graph rather than regenerating it.
  • Entity disambiguation. Mentions of the same entity across a document are merged into one.
  • Validation. Deterministic checks, SHACL shapes, and repairs that need no extra model call.
  • Provenance. Facts carry RDF 1.2 provenance back to the text, which you can strip on output.
  • Ontology context. Each part of a document is shown the ontology it needs: chosen from your catalog, retrieved from a vector store (LanceDB or Qdrant), or fixed.
  • Storage. In memory by default, Apache Jena Fuseki for persistence, partitioned by tenant and project.
  • A light core. The base install embeds without a document-processing stack or an ML runtime; those are extras.

Install

pip install "ontocast[server,openai,doc-processing]"

server provides the ontocast command and HTTP API, openai the model provider (anthropic, google and ollama also exist), and doc-processing the document converter (PDF, Office, HTML, Markdown, images). All extras: Installation.

Quick start

OntoCast reads its settings from environment variables:

export LLM_API_KEY=sk-...
ontocast serve
curl -X POST http://127.0.0.1:8999/process -F "file=@document.pdf" -o result.json

The response holds the facts and the ontology as Turtle. To process files without a server:

ontocast process --input-path ./papers --output-dir ./out

To keep settings in a file, copy .env.example.minimal to .env and load it into your shell with set -a; source .env; set +a: OntoCast does not read the file itself. Step by step: Quick start.

Your own ontologies

Put Turtle files in a directory and pass it at startup, or upload them to a running server:

ontocast serve --ontology-dir ./my-ontologies
curl -X POST http://127.0.0.1:8999/ontologies -F "file=@my-ontology.ttl"

Embed in your agent

from langchain.agents import create_agent
from ontocast import Config, ToolBox, ontocast_tools

tools = await ToolBox.acreate(Config.in_memory())
await tools.initialize()

agent = create_agent(
    model,
    tools=[*ontocast_tools(tools)],
    prompt="Edit the ontology from the user's text.",
)

run_unit_pipeline processes a single passage, and make_ontocast_node adds OntoCast to your own LangGraph. See Embedding OntoCast.

How it works

The OntoCast pipeline: convert, chunk, then the ontology stages and the facts stages, then serialize

A document is converted to text and cut into parts. Each part updates the ontology; the updates are normalized, consolidated and checked. Each part then yields facts in the ontology's terms; the facts are merged, disambiguated and validated, and the result is written to the triple store. How OntoCast works.

Documentation

Getting started Installation · Quick start
Concepts How OntoCast works · Ontologies and facts
Guides Configuring OntoCast · Recipes · Validation and SHACL · Triple stores
Reference Configuration · HTTP API · Python API

Release notes: CHANGELOG.md

Contributing

See Contributing. Issues and discussion: GitHub. Contributors accept the Contributor License Agreement once, by commenting on their first pull request.

License

Apache License 2.0; see LICENSE.

Metadata

Release files for ontocast 0.6.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ontocast 0.6.5
File Size Uploaded
ontocast-0.6.5.tar.gz 1.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for ontocast 0.6.5
File Interpreter ABI Platform
ontocast-0.6.5-py3-none-any.whl Python 3 none any Details

Total release size: 2.0 MB

Release files / ontocast-0.6.5.tar.gz

Download URL ontocast-0.6.5.tar.gz
Size 1.2 MB
Tags Source
SHA-256 checksum
How to use checksums
4812bfbb52030b10f289bbad0af8de786a9b83e5fbc5b8abf661de828daccfc5
BLAKE2b-256 checksum
How to use checksums
747f911dd663be548f4387dad91682650cec81dbe0c988efccc1ac602628de55
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / ontocast-0.6.5-py3-none-any.whl

Download URL ontocast-0.6.5-py3-none-any.whl
Size 799.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d6404b6d737f6812bef6ae081672ccc7f17e530b3f5d806509542604b4b4f8c4
BLAKE2b-256 checksum
How to use checksums
88863aee7d57150432b0cbb4c43f79a67430269b4073db6f9486b3f982dcdb56
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.6.7

2 release files

0.6.6

2 release files

This release

0.6.5 This release

2 release files

0.6.4

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.4.3

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page