nmdc-metadata-suggestor-ai-tool
A Python application for the NMDC Submission portal metadata suggestor tool, powered by AI. This project uses modern Python tooling with uv for dependency management.
Prerequisites
- Python 3.12 or higher
- uv (or use Docker)
- Docker and Docker Compose (for containerized development)
Quick Start
LLM Configuration:
You will need to set up a .env file. Copy the example first:
cp .env.example .env
Environment variables used by LLMClient and ConversationManager:
AI_INCUBATOR_KEY: API key for PNNL AI Incubator (when usingaccess_provider=pnnl).AI_INCUBATOR_BASE_URL: Base URL for the PNNL AI Incubator API.GOOGLE_APPLICATION_CREDENTIALS: Path to a GCP service account JSON file (for Vertex AI).VERTEX_PROJECT_ID: (Optional) GCP project id for Vertex. If not provided, the SDK will attempt to infer it from credentials.GCP_REGION: (Optional) Vertex region override for Gemini calls (falls back toGCP_REGION, thenus-east5).CLAUDE_CODE_USE_VERTEX(Required by ConversationManager.agentic()) : Forcesclaude_agent_sdkto use vertex AI for credentials. Required for our use as we use GCP for agentic auth.CBORG_KEY: API key for CBORG (when usingaccess_provider=cborg).CBORG_BASE_URL: Base URL for the CBORG API.
The LLMClient will read the appropriate variables depending on access_provider (set to pnnl, cborg, or gcp).
Environment variables are loaded from a .env file in the project root via
python-dotenv. Variables already
set in your shell take precedence over .env values (override=False is the
default).
Agent permissions
ConversationManager.agentic() reads .claude/settings.json for its tool allowlist. Without it
the headless agent cannot run ontology lookups and answers from its prompt alone — silently.
See docs/agent-permissions.md.
ENVO ontology cache
Env triad suggestions resolve ENVO terms through oaklib.
On first use oaklib downloads the ENVO semantic-sql build (~15 MB) and caches it under ~/.data/oaklib;
later runs read the cache. Warm it ahead of time — in a container image, or before a first run offline:
uv run python -c "from oaklib import get_adapter; get_adapter('sqlite:obo:envo')"
Set ENVO_ADAPTER to point at a different build (a pinned or self-hosted one) if you need to.
Using uv (Local Development)
-
Install uv (if not already installed):
curl -LsSf https://astral.sh/uv/install.sh | sh # or pip install uv
-
Clone and setup:
git clone https://github.com/microbiomedata/nmdc-metadata-suggestor-ai-tool.git cd nmdc-metadata-suggestor-ai-tool
-
Install dependencies:
uv sync -
Configure environment:
cp .env.example .env # Edit .env and add your API keys
-
Use the package in Python:
uv run python
import json from nmdc_metadata_suggestor_ai_tool.llm_client import LLMClient from nmdc_metadata_suggestor_ai_tool.recommendation_pipeline import run_recommendation_pipeline # Any NMDC submission JSON payload. This example fixture ships with the repo: # a soil study with a data DOI, so the run also exercises DOI abstract ingestion. with open("tests/fixtures/test_submission.json") as f: submission_object = json.load(f) client = LLMClient(access_provider="gcp") result = run_recommendation_pipeline(submission_object, client) print(result.model_dump_json(indent=2))
Each suggestion names an NMDC field the submitter should fill in and a
reasonciting the submission text that supports it.valueis filled in only when the submission contains that value literally; when the value is inferred, the inference goes in thereasonandvaluestays"". An emptyvalueis expected output, not a failure. (The env triad path is the exception — it always returns alabel [CURIE]value.)
Advanced: direct ConversationManager usage (optional)
from nmdc_metadata_suggestor_ai_tool.llm_client import LLMClient, ConversationManager
from nmdc_metadata_suggestor_ai_tool.system_prompt import system_prompt
client = LLMClient(access_provider="gcp")
conversation = ConversationManager(llm_client=client, system_prompt=system_prompt)
# Add plain text context (pdf_files may be a list of local PDF paths)
conversation.add_message(text="Please summarize the submission.", pdf_files=None)
# Add any schema context to guide the model
conversation.add_schema_context("<schema description here>")
response = conversation.generate(model="gemini-2.5-flash", max_tokens=1024, gemini_temperature=0.2)
print(response)
Langfuse set up
This project uses Langfuse to track LLM logging. We have a pro plan using the cloud-hosted Langfuse. Configuration is simple:
- Obtain your API keys from the Langfuse UI under Settings → API Keys (US cloud: https://us.cloud.langfuse.com).
- Copy
.env.exampleto.envand fill in the Langfuse section:
LANGFUSE_PUBLIC_KEY=pk-lf-...
LANGFUSE_SECRET_KEY=sk-lf-...
LANGFUSE_BASE_URL=https://us.cloud.langfuse.com
LANGFUSE_TRACING_ENVIRONMENT=local # options: production, development, local, unknown
Leaving these variables unset disables tracing entirely.
Development
Running Tests
# Run all tests
uv run pytest
# Run with coverage
uv run pytest --cov=src/nmdc_metadata_suggestor_ai_tool
# Run specific test file
uv run pytest tests/test_recommendation_pipeline.py
Code Quality
# Format code with Ruff
uv run ruff format
# Lint with Ruff
uv run ruff check
# Type check with MyPy
uv run mypy src
Adding Dependencies
# Add a production dependency
uv add package-name
# Add a development dependency
uv add --dev package-name
# Update dependencies
uv sync
Configuration
Configuration is managed through environment variables or a .env file. See .env.example for all options. Beyond the LLM access variables above:
CONTACT_EMAIL: Contact email sent in User-Agent headers for the Crossref polite pool and OpenAlex (defaults tosupport@microbiomedata.org).
Advanced ingestion tuning (all optional; defaults shown):
NMDC_EDI_MAX_XML_CHARS: Max characters read from an untrusted EDI metadata XML payload (default2000000).NMDC_DATAONE_SOLR_MAX_XML_CHARS: Max characters read from an untrusted DataONE Solr XML payload (default2000000).NMDC_EUROPEPMC_MAX_XML_CHARS: Max characters read from an untrusted Europe PMC full text XML payload (default2000000).NMDC_HTTP_RETRY_ATTEMPTS: Retry attempts for publication/DOI HTTP requests (default3).NMDC_HTTP_RETRY_BACKOFF_SECONDS: Base backoff in seconds between retries (default0).NMDC_HTTP_MAX_RETRY_DELAY_SECONDS: Cap in seconds on retry delay (default30).NMDC_HTTP_POOL_CONNECTIONS: HTTP connection pool size (default20).NMDC_HTTP_POOL_MAXSIZE: HTTP connection pool max size (default100).
Contributing
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Make your changes
- Run tests and quality checks
- Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
License
See LICENSE for licensing terms.
Metadata
Release files for nmdc-metadata-suggestor-ai-tool 1.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| nmdc_metadata_suggestor_ai_tool-1.3.0.tar.gz | 210.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| nmdc_metadata_suggestor_ai_tool-1.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 350.3 kB
Release files / nmdc_metadata_suggestor_ai_tool-1.3.0.tar.gz
| Download URL | nmdc_metadata_suggestor_ai_tool-1.3.0.tar.gz |
|---|---|
| Size | 210.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
709ac0e82d29512ecabcb06ab9c66092c15f1e0af632a6323df009463bede4e8
|
|
BLAKE2b-256 checksum How to use checksums |
279d8bb8e6be1b44f748ed4c0f02960044989d783d18d697fba6521d8b16bf90
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency logRelease files / nmdc_metadata_suggestor_ai_tool-1.3.0-py3-none-any.whl
| Download URL | nmdc_metadata_suggestor_ai_tool-1.3.0-py3-none-any.whl |
|---|---|
| Size | 140.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
93b553a2cb08deef37596d77ef9d1d1f7c253856e9991ad570aef0e306509115
|
|
BLAKE2b-256 checksum How to use checksums |
259495585dc44187bfc7586a89c0839cda11e19f448402b413dbe0d76865990c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency log