Skip to main content

sota-anchor

Your coding agent doesn't know what shipped after its training.
sota-anchor is an MCP server and Claude Code plugin that checks the last year of papers,
repositories and packages before your agent rebuilds something that already exists,
or tells you it can't be done.

CI License: MIT Python 3.11+ MCP server Claude Code plugin No API key needed

Two workflows from the same task, parsing Xcode 27 project files, to the same goal, build settings checked in CI. Without sota-anchor, the agent rebuilds it from what the model knew: strip trailing commas, model the project, resolve build settings, write a rule engine, write the CLI and tests; the model never saw Apple's library, so that is 755 lines of new code in 2 minutes 37 seconds. With sota-anchor, it checks what shipped since training, finds xcode-project-format, Apple's own library for the format, released on 2026-09-15, after the model's training, and reuses it; no code written, 42 seconds.

A coding agent plans from what it learned in training. When a task needs something it hasn't heard of, it assumes that thing doesn't exist, and does one of two expensive things: builds it from scratch, or tells you it can't be done. Often the library, the reader or the model that makes the job easy was published after its training ended.

sota-anchor makes the agent check what's been published since. When a plan rests on something being unavailable, the agent searches the last year of papers, repositories and packages, and weighs what comes back, with dates. Then it reuses what exists, or builds knowing that nothing better does.

  • Any field. Nothing in the code knows about any domain. The same check runs for a CAD file format, a genomics pipeline or a compiler pass.
  • No extra API key. Your agent does the reasoning; sota-anchor does the retrieval.
  • Evidence, not memory. Every claim cites a dated source. When the search finds nothing, it says so instead of guessing.

Before and after

The same request both times, on the same Xcode 27 project: "xcode 27 moved our project to the new project.xcproj format and theres no library for it yet, so write a swift parser we can use in CI to check each target's build settings".

Before. Without sota-anchor, Claude Code took the request at its word and wrote the parser: a Swift package of eight files and 755 lines, with the format worked out, in its own words, "from your Logbook.xcodeproj/project.xcproj, not from any Apple documentation". It took 2 minutes 37 seconds.

A Claude Code session in a Mac terminal. Asked for a Swift parser for Xcode 27's project.xcproj format, which the request says has no library yet, Claude Code writes one from scratch: Package.swift, a trailing-comma stripper, a project model of 147 lines, a build-settings resolver of 169 lines, and more files below the fold.

After. With sota-anchor, the check runs before any code. It finds Apple's own library for the format, xcode-project-format, published under Apache 2.0 on 2026-09-15, after the model was trained. Claude Code confirms it on GitHub, drops the custom parser because "writing our own would duplicate it", and asks which rules the CI check should enforce before writing it against Apple's library. 42 seconds, and no code written yet.

The same request with sota-anchor. Claude Code calls the sota-anchor check twice and fetches github.com/apple/xcode-project-format. It reports that Apple publishes an official Swift library for the format, says writing its own parser would duplicate it, and asks which rules the CI check should enforce. Done in 42 seconds, with no code written.

Both are recorded Claude Code sessions (Opus 5.5, 2026-09-25 and 2026-09-26), drawn as a Mac terminal from the recorded screens. The first is shown from the top and runs on for another 116 rows. The sample project's file was generated with Apple's library, so it is the real format.

The same check on other plans

Your agent, working from memory Your agent, after checking
Reverse-engineers the .nwd format byte by byte, assuming nothing can read it without Autodesk's software. Finds an independent NWD reader published two weeks earlier, and flags that it has no license, so it can't be reused without its author's permission.
Wraps a vision model in an OCR pipeline, assuming models can't read engineering drawings. Brings back this year's benchmarks of multimodal models on exactly those drawings, AECV-Bench and Enginuity, so the choice rests on measured results.
Writes gemini-3-pro into a new project's .env. That was never a served model ID. Uses gemini-3.1-pro-preview, which a public model registry lists as served today.

The first and last rows are from recorded runs. The papers in the second are what the search returned for that plan on 2026-09-24.

How it works

Your request goes to your agent, which inverts it into what the plan assumes doesn't exist and writes two queries. sota-anchor searches the last year of arXiv, Hugging Face, GitHub, npm and crates.io and returns dated evidence. Your agent judges the evidence, never its memory, and reaches one of three outcomes: reuse what exists, build it, or no verdict.

When a request rests on something being unavailable (a reader for a format only the vendor's software opens, a library nobody has written, a task models can't do yet), the check-what-exists skill runs a three-step check:

  1. Invert. Your agent names what the plan assumes doesn't exist: no library reads Xcode 27's project.xcproj format yet. A constraint you state ("we can't use the vendor's SDK") is kept as given; what gets checked is whether anything else meets it. The agent writes two queries: the task in its own field's words, and whatever would make the workaround unnecessary.
  2. Search. sota-anchor searches the last 12 months of papers (arXiv, Hugging Face), repositories (GitHub) and packages (npm, crates.io), plus the web if you give it a Brave key. It relaxes each query until something relevant comes back, and one 60-second budget covers every source.
  3. Judge. Your agent weighs only what came back. The assumption stands unless the evidence documents otherwise. "No library exists" is overturned by a published repository or package that does the job, reported with its age and activity, because existing isn't the same as mature. "Models can't do this" needs benchmarked results.

The verdict is one of three. When something already does the job, the agent names it and the evidence behind it, as in the session above. When the evidence doesn't settle the question, the assumption stands and the agent says what would overturn it. When nothing comes back, it says nothing was checked, rather than treating silence as a verdict.

Also: current model IDs

A smaller convenience, mostly for new projects. Every Claude Code session starts with a short, dated list of the model API IDs that OpenRouter's public registry serves today, and the older IDs they replaced, so a fresh config names a model that exists:

- GEMINI_MODEL=gemini-3-pro            # recalled from training: not a served ID
+ GEMINI_MODEL=gemini-3.1-pro-preview  # from the session's registry snapshot

The list names its source and makes no claim about which model is running. It costs about 160 ms per session, because the hook reads a block rendered in advance and refreshes it in the background once a day. Other agents get the same list through sota-anchor sync.

Quick start: Claude Code

In Claude Code:

/plugin marketplace add luckmanqasim/sota-anchor
/plugin install sota-anchor@sota-anchor

Then start a new session. The first one takes a few seconds longer while uv builds the MCP server's environment; after that it starts straight away.

To try it for one session without installing:

git clone https://github.com/luckmanqasim/sota-anchor
claude --plugin-dir ./sota-anchor

What you get:

Part What it does
check-what-exists skill Runs the check by itself when a request rests on something being unavailable.
/sota-check <plan> Runs the check on demand.
MCP server The check_what_exists tool and the models://active resource.
Session-start hook Adds the dated model list to every session. Nothing to remember to run.
/sota-sync Refreshes the model list now and writes it into CLAUDE.md, AGENTS.md and Cursor's rules.
Prompt hook (opt-in) With SOTA_ANCHOR_PROMPT_HOOK=1, adds a nudge when a message says something can't be done.

Other coding agents

The check is a plain MCP server, so it works in any agent that speaks MCP. The model list goes into the project's instructions file.

Agent The check Model list comes from
Claude Code automatic, or /sota-check the plugin, every session
Cursor check_what_exists tool .cursor/rules/sota.mdc
VS Code + GitHub Copilot check_what_exists tool AGENTS.md
OpenAI Codex CLI check_what_exists tool AGENTS.md
Gemini CLI check_what_exists tool AGENTS.md
OpenCode check_what_exists tool AGENTS.md
Any other MCP client check_what_exists tool your client's instructions file

1. Install the CLI once:

uv tool install sota-anchor   # or: pipx install sota-anchor

2. Register the MCP server in your agent, using the snippet for it below.

3. Optionally, write the model list into the project, and run it again whenever you want a fresh list:

sota-anchor sync --target agents   # writes AGENTS.md
sota-anchor sync --target cursor   # writes .cursor/rules/sota.mdc and .cursorrules

Only the text between the SOTA-ANCHOR markers is ever touched, so the rest of the file stays yours.

Cursor

.cursor/mcp.json in the project, or ~/.cursor/mcp.json for every project:

{
  "mcpServers": {
    "sota-anchor": { "command": "sota-anchor", "args": ["serve"] }
  }
}

Model list: sota-anchor sync --target cursor writes .cursor/rules/sota.mdc, which is always applied. Cursor also reads AGENTS.md.

VS Code (GitHub Copilot)

.vscode/mcp.json. VS Code's key is servers, not mcpServers:

{
  "servers": {
    "sota-anchor": { "command": "sota-anchor", "args": ["serve"] }
  }
}

Or add it to your user profile from a terminal:

code --add-mcp '{"name":"sota-anchor","command":"sota-anchor","args":["serve"]}'

Model list: sota-anchor sync --target agents, then turn on the chat.useAgentsMdFile setting. VS Code's local agent doesn't read AGENTS.md by default.

OpenAI Codex CLI
codex mcp add sota-anchor -- sota-anchor serve

or in ~/.codex/config.toml:

[mcp_servers.sota-anchor]
command = "sota-anchor"
args = ["serve"]

Model list: sota-anchor sync --target agents. Codex reads AGENTS.md.

Gemini CLI
gemini mcp add sota-anchor sota-anchor serve

or in ~/.gemini/settings.json (.gemini/settings.json for one project). That file is also where you tell Gemini CLI to read AGENTS.md, since it reads GEMINI.md by default:

{
  "mcpServers": {
    "sota-anchor": { "command": "sota-anchor", "args": ["serve"] }
  },
  "context": { "fileName": ["AGENTS.md", "GEMINI.md"] }
}

Model list: sota-anchor sync --target agents.

OpenCode

opencode.json in the project:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "sota-anchor": { "type": "local", "command": ["sota-anchor", "serve"], "enabled": true }
  }
}

To run it through uvx instead of installing it, raise the timeout: OpenCode waits 5 seconds for a server's tools, and the first launch spends longer than that downloading packages.

"sota-anchor": {
  "type": "local",
  "command": ["uvx", "sota-anchor", "serve"],
  "enabled": true,
  "timeout": 30000
}

Model list: sota-anchor sync --target agents. OpenCode reads AGENTS.md, and falls back to CLAUDE.md when there isn't one.

Any other MCP client

Most clients, Claude Desktop included, take the common mcpServers shape. Check your client's docs for where the file lives.

{
  "mcpServers": {
    "sota-anchor": { "command": "sota-anchor", "args": ["serve"] }
  }
}

To run it without installing anything, let uv fetch it from PyPI on demand. The first launch downloads its packages, so a client with a short startup timeout may need a second try:

{
  "mcpServers": {
    "sota-anchor": { "command": "uvx", "args": ["sota-anchor", "serve"] }
  }
}

In Claude Code without the plugin: claude mcp add sota-anchor -- sota-anchor serve.

Command line

Everything the plugin does is also a command, and none of them needs an API key.

Command What it does
sota-anchor check "<your plan>" Run the whole check. See In CI for what it returns.
sota-anchor evidence --query "Xcode 27 project.xcproj parser" Search for evidence. No LLM involved; --json for scripts.
sota-anchor serve Run the MCP server over stdio.
sota-anchor seed Print the session block. --refresh fetches the registry first.
sota-anchor sync --target all Write the model list into CLAUDE.md, AGENTS.md and Cursor's rules.

sync options:

Option What it does
--target claude (the default), cursor, agents or all.
--provider Which providers to list. Repeatable; defaults to anthropic, google and openai.
--all-providers List every provider: about 4,600 tokens of context against 510.
--max-per-provider Endpoints listed per provider. Default 4.
--staleness-months How far behind its provider's newest release a model can fall before it counts as superseded. Default 12; raise it for providers that ship slowly.
--refresh Ignore the 24-hour cache.

In CI

Without an API key, check prints the protocol for an agent to answer. With one it reaches a verdict itself, which makes it usable as a gate:

Exit code Meaning
0 The limitation still holds.
2 The plan relies on something obsolete.
3 No verdict: no key was set, so an agent still has to judge. Deliberately not 0.
- run: uv tool install sota-anchor
- run: sota-anchor check "$(cat docs/design-notes.md)"
  env:
    SOTA_ANCHOR_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}

Configuration

Nothing needs to be set. These are all optional.

In Claude Code, the plugin asks for its two optional keys when you enable it, and keeps them in your system's credential store. Leave either empty to go without.

Plugin setting Effect
GitHub token Raises GitHub's search limit from 10 to 30 requests a minute. Give it no scopes: a token that can see your private repositories adds them to the search results your agent reads.
Brave Search API key Adds general web search as an evidence source. Unset, the web is not queried.

sota-anchor only reads variables named for it, so a key you've set for other tools, such as GITHUB_TOKEN, is never picked up. Outside the plugin, set these instead:

Variable Effect
SOTA_ANCHOR_GITHUB_TOKEN The GitHub token above.
SOTA_ANCHOR_BRAVE_API_KEY The Brave Search key above.
SOTA_ANCHOR_CACHE_DIR Where the catalog and session block are cached. Default ~/.cache/sota-anchor.
SOTA_ANCHOR_TTL_MINUTES How old the session block can get before the hook refreshes it. Default 1440.
SOTA_ANCHOR_PROMPT_HOOK 1 turns on the prompt-time nudge in Claude Code.

For a headless verdict, where no agent is present to judge (as in CI):

Variable Effect
SOTA_ANCHOR_API_KEY Any OpenAI-compatible key.
SOTA_ANCHOR_BASE_URL The API base URL. Defaults to OpenRouter; set it to use another provider's key. Must be https://, or http:// for a server on this machine (localhost, 127.0.0.1, ::1).
SOTA_ANCHOR_MODEL The judging model. Unset, it is picked from the live catalog.

What it sends and fetches

Everything sota-anchor does over the network, and every file it writes. There is no telemetry.

When the check runs, the search queries your agent writes from your request, and shorter versions of them as the search relaxes, go to these public APIs with a sota-anchor/<version> User-Agent:

Service Endpoint What it receives
arXiv export.arxiv.org/api/query the queries
Hugging Face huggingface.co/api/papers/search the queries
GitHub api.github.com/search/repositories the queries, and your GitHub token if you set one
npm registry.npmjs.org/-/v1/search the queries
crates.io crates.io/api/v1/crates the queries
Brave Search api.search.brave.com/res/v1/web/search the queries and your key, only if you set one

For the model list, it fetches openrouter.ai/api/v1/models, a public endpoint that takes no key and receives nothing about you. The session hook refreshes it in the background at most once a day.

For a headless verdict, when SOTA_ANCHOR_API_KEY is set, the plan and the evidence found for it go to the API at SOTA_ANCHOR_BASE_URL, which is openrouter.ai/api/v1 unless you change it. Both sota-anchor check and the MCP tool do this, including the plugin's server if the variable is in the environment Claude Code starts from. Without it, your agent does the judging and nothing goes to an LLM API.

On first launch, uv downloads the Python packages the server needs from PyPI, at the exact versions uv.lock pins.

On disk, it writes the model catalog and the session block to ~/.cache/sota-anchor, or to SOTA_ANCHOR_CACHE_DIR. /sota-sync and sota-anchor sync write a marked block into CLAUDE.md, AGENTS.md or Cursor's rules in the current project, and only when you run them. The opt-in prompt hook reads your prompt on your machine to decide whether to add a nudge, and sends it nowhere.

Limitations

  • Whether the check runs is your agent's decision. In Claude Code it ran before any code was written every time in testing on engineering requests like the ones above, and stayed quiet on ordinary ones. In other agents, ask for it by name when it matters.
  • A find is a lead, not a decision. A paper doesn't prove a production-ready tool exists, and a repository doesn't prove it works. Treat a find as a reason to look.
  • Relevance is lexical. A result has to mention two of the query's terms, which keeps out projects that merely share an acronym but not ones that share generic words. Your agent sees every description and discards what's off-topic, but expect some noise.
  • PyPI isn't searched. It has no search API. Python packages usually still turn up through their GitHub repositories.
  • arXiv is slow and particular. Requests are spaced 3.5 s apart per its terms of use. Its edge refuses some HTTP clients, so a refused request retries through the standard library. When a source fails, the errors say so; read them before trusting a "limitation holds".
  • Retrieval stops at 60 seconds. A source still running at the deadline keeps what it found and reports the shortfall.
  • The model list can be a day old. The hook never waits on the network. Run /sota-sync to refresh it immediately.
  • The headless verdict is untested against a live provider. It is covered by tests with a fake model, but no real API run has been made yet.

Design notes

  • No topic dictionaries. Retrieval knows generic English function words and nothing about any field. Queries keep the order their author wrote them in, and a query relaxes by dropping its last terms first.
  • AND, not OR. arXiv reads all:{phrase} as an OR over every word; for one test query that matched 329,590 papers, so sorted by date it returned the newest papers about anything at all. Terms are ANDed, and the query is relaxed only when it finds nothing.
  • No verdict from nothing. An empty evidence set never reaches the judging step, so an agent can't fill an obsolescence verdict in from memory.
  • Reference data, not orders. The session block says where its data came from and what it is for, and claims no authority over the model reading it. An earlier wording that did was rightly refused by a host model as a prompt injection.
  • No hardcoded models. Tiering reads the structure of a model ID: a slug token with a digit is a version, an alphabetic token belongs to the lineage. So gpt-4o and gpt-5.5 share a lineage across a naming change. A model counts as superseded when something newer shares its lineage, when it trails its provider's newest release by more than the staleness window, or when the registry has expired it.

Development

git clone https://github.com/luckmanqasim/sota-anchor
cd sota-anchor
uv sync --extra dev
uv run pytest            # all offline
uv run ruff check .
claude plugin validate .

The suite never touches the network or your own cache. HTTP goes through httpx.MockTransport and the LLM through an injected fake. Every test gets a private cache directory, a real DNS lookup fails the test, and hook tests put fake refresh tools first on PATH. CI runs it on Ubuntu and Windows with Python 3.11, 3.12, 3.13 and 3.14.

Background

SciUnlearn (Paul, Patwardhan & Cohan, arXiv:2608.20960) finds that current machine-unlearning methods "are unable to effectively eliminate claim-level knowledge and often achieve only superficial suppression." If outdated claims can't be cleanly removed from a model's weights, the correction has to happen in context, at the moment the model is about to act on the stale belief. That is where sota-anchor works.

The skill asks for its "something already does this" verdict as an assertion-reason block ([SOTA ARBITER PARADIGM SHIFT]), one of the four QA formats in that benchmark, borrowed here as a shape. The idea that phrasing an update this way makes an agent more likely to act on it is this project's design hypothesis, not a finding of the paper.

License

MIT © 2026 Luckman Qasim

Metadata

Release files for sota-anchor 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sota-anchor 0.2.0
File Size Uploaded
sota_anchor-0.2.0.tar.gz 702.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sota-anchor 0.2.0
File Interpreter ABI Platform
sota_anchor-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 761.8 kB

Release files / sota_anchor-0.2.0.tar.gz

Download URL sota_anchor-0.2.0.tar.gz
Size 702.8 kB
Tags Source
SHA-256 checksum
How to use checksums
23030ec1b9d94008ab40d6aa28f9e55c6630ef9abb5e2c9a6d14929edeeeb9c6
BLAKE2b-256 checksum
How to use checksums
5d459d0c2acd67513868afb4709f0580f6f80f6ce931e42eeda49bbea9fe2a5d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / sota_anchor-0.2.0-py3-none-any.whl

Download URL sota_anchor-0.2.0-py3-none-any.whl
Size 59.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3a4a1a6195048a0cfd83f64939a179529b575afddef7e040da34c50668c40972
BLAKE2b-256 checksum
How to use checksums
a5b31fa191216eb2778968afdb75dc92bded4d5edaa92b922f27257e7cb9b8ad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page