Skip to main content

Persistent structural context and ultra-fast repeated analysis for AI coding agents

Project description

ASK Engine

ASK — Actionable Software Knowledge. Persistent structural intelligence for AI coding agents.

Context · Impact · Migration · Architecture · Review — everything from one structural model.

Version Python

ASK Engine is the product. The CLI command is ask. The legacy sourcecode command still works as a deprecated alias (it prints a one-line notice and forwards to ask) and remains the Python/PyPI package name for now. The authoritative version is whatever ask version reports. See docs/PRODUCT_IDENTITY.md.


The problem

Every time an AI coding agent starts a new session, it has to re-parse the repository from scratch. For a large Java or TypeScript monolith, that means 5–15 seconds per invocation. Multiply by dozens of agent turns per hour, and repo context acquisition becomes a real bottleneck — not just latency, but tokens, compute, and iteration velocity.

ASK Engine solves this with a persistent structural cache keyed on file content hashes. After the first scan, every subsequent invocation returns pre-built context in milliseconds. The repo doesn't change? The cache doesn't expire.

The cache is not a performance optimization. It is what makes ASK Engine usable as infrastructure rather than a one-off tool.


Proof — measured on real repos

Repo Size Cold scan Cache hit Speedup
Keycloak 7,885 Java files 10.5s 0.6s ~17x
BroadleafCommerce 2,985 Java files 2.7s 0.3s ~9x

Cache keyed on content hashes — invalidated only when source changes. On repeated agent sessions against the same codebase, nearly every invocation is a cache hit.

At 0.3s per call, ASK Engine becomes constant infrastructure inside agent loops — call it before every edit, every PR review, every test run, without batching or caching manually.

What a warm actually covers. ask cache warm runs the compact analysis: it rebuilds the shared structural layers (L1/L2 + the Repository Intelligence Snapshot + the shared Canonical IR) and the compact view. Pass --agent to warm the agent view as well. Deeper projections are separate keys and are not covered by either — --full, --env-map and a raised --depth recompute on first use, as do most prepare-context tasks. Measured on a 3,342-file Spring monolith: cache warm 103s → --compact --git-context 1s (hit), but --agent --full --env-map --depth 20 still 171s (miss).

One rule for invalidation, and a table that names every command. Every layer keys on the exact tree state: any change to the analysed files invalidates it, committed or not. Which of your commands a warm helps — the answer, only the shared work, or nothing — is published per command: ask cache model, or docs/CACHE.md.


Install

# Homebrew (macOS / Linux)
brew tap haroundominique/sourcecode && brew install sourcecode

# pip / pipx
pipx install sourcecode        # or: pip install sourcecode

ask version                    # ask 2.5.1

Package vs. command. The install package is named sourcecode this release (renaming the distribution is a separate, breaking change). Installing it gives you the canonical ask command plus the deprecated sourcecode alias.


Quickstart

Start with these four. In three independent field evaluations they carried most of the measured value, and posture --diff is the one no evaluator found an equivalent for — commercial or open source.

# What does each profile set ACTUALLY wire — and what changes between them?
# Resolves conditional beans and the filter chain, then diffs effective endpoint access.
ask posture /path/to/repo --diff dev:prod

# Every REST endpoint with its effective path (context-path + servlet path resolved),
# its inferred security policy and a confidence per endpoint.
ask endpoints /path/to/repo

# Spring semantic audit: transactional anomalies (private @Transactional = silent
# CGLIB no-op), security surface, request-body validation.
ask spring-audit /path/to/repo

# Spring Boot 2→3 readiness: located blockers, per-dimension score, effort estimate.
ask migrate-check /path/to/repo --compact

Then the three that finish the sentence the four above start. A profile-set answer stays conditional until you know which set the deployment starts with, a gate nobody can populate is a gate nobody runs, and a defect ranked without its reach is a linter note:

# Which profile set ACTUALLY runs? Reads every artefact that can set it — build,
# descriptors, Dockerfile, Compose, Kubernetes, launch scripts — with file:line, and
# says NOT DECIDED IN THIS REPOSITORY when nothing in the tree decides it. Then ranks
# every admissible set by what it leaves reachable without authentication.
ask posture /path/to/repo --resolve-environments

# Derive the contracts this repository already satisfies, execute each one before
# writing it, and populate .ask/contracts.yml from the measurement instead of by hand.
ask verify /path/to/repo --init

# What each defect actually costs once reach, access and write effect are in it:
# severity × reachability × auth_verdict × write_effect, ordered, every factor traceable.
ask risk /path/to/repo

Then the everyday loop:

# High-signal structural summary — warm cache ~0.3s, cold 2–10s
ask --compact

# Blast radius: what breaks if this class changes?  (target the INTERFACE, not the Impl)
ask impact OrderService /path/to/repo

# Onboard to an unfamiliar codebase
ask onboard /path/to/repo

# PR review: risk, test gaps, changed modules
ask review-pr /path/to/repo --since main

# CI gate on NEW violations only, instead of on pre-existing debt
ask verify /path/to/repo --init               # derive the contracts, don't hand-write them
ask verify /path/to/repo --capture-baseline   # accept today's debt, once
ask verify /path/to/repo                      # then: only new violations block

Adopting a gate on a real codebase. Nobody hand-writes contracts for a 3 000-file monolith, so ask verify --init derives the ones the repository satisfies today and executes each before writing it — init followed by verify passes by construction. A repository that starts declaring contracts already violates them somewhere; ask verify is baseline-relative by default (--fail-on new) so the gate survives contact with reality instead of being switched off on day one. ask baseline capture|diff|trend is a different thing: versioned architectural metrics over time, for trend reporting rather than blocking.

Full command reference: docs/USER_GUIDE.md · posture in depth: docs/posture.md.


Capabilities

Everything is computed from one cached structural model. Seven groups:

1 · Structural Context

Bounded, noise-free repo context designed to drop straight into an agent's context window. ask --compact · ask --agent · ask onboard · ask cold-startreference

2 · Impact Analysis

Blast radius from a class or interface — reverse dependencies, through Spring DI, to the HTTP endpoints a change reaches. ask impact · ask impact-chain (TX/SEC-enriched) · ask pr-impactreference

3 · Architecture Intelligence

The system map: module graph, dependency views, REST surface, per-class summaries, and a symbol-level IR for downstream tooling. ask export · ask repo-ir · ask endpoints · ask explainreference

4 · Migration & Modernization

Is this codebase ready to upgrade? Per-dimension readiness (Jakarta / Spring Boot / JDK / Hibernate), located blockers, an effort estimate, and — with --blast-radius — the endpoints whose call path runs through each blocker, so the re-test plan is ordered by regression scope. ask migrate-check · ask modernizemigrate-check reference · MODERNIZATION.md

5 · Spring Analysis

Deterministic Spring semantics: transactional anomalies (e.g. @Transactional on a private method = silent CGLIB no-op), security surface, request-body validation. ask spring-audit · ask validationreference

5b · Runtime Posture (experimental — and the most differentiated thing here)

What a profile set actually wires: which conditional beans register, which do not, and which conditions could not be decided at all — then the effective endpoint access that follows from the filter chain. --diff answers the question nobody else answers in one command: what changes between dev and prod, across every endpoint at once. ask posture · ask posture --diff dev:prod · ask posture --property k=vposture.md

--resolve-environments carries it past the edge of the repository. Every answer above is conditional on a profile set, and which set runs is decided by the deployment. This flag reads every artefact in the tree that can set spring.profiles.active — Maven/Gradle, web.xml, Dockerfile, Compose, Kubernetes, launch scripts — with file:line, and returns decided_in_repository, artefacts_disagree (two deployments, never resolved by picking one) or NOT DECIDED IN THIS REPOSITORY. Then it resolves every admissible set over one parse and publishes the worst. A field audit closed those five hops by hand with grep before the flag existed. ask posture . --resolve-environments

Unresolved is a first-class outcome: a condition the resolver cannot decide is reported as a hole with the condition named, never folded into active or inactive. A posture answer that guesses is a confident security falsehood — the worst failure mode this tool has.

5c · Composed Risk (experimental)

Every command above answers one axis, and a reader composes them by hand. ask risk does the join over the endpoint and symbol ids these commands already share with each other: severity_effective = defect_severity × reachability × auth_verdict × write_effect. The product of the published factors is the published score — decomposable to the four figures and the authority behind each — and an axis this build cannot measure is unknown, weighted 1.0 and named in blind_axes rather than silently treated as safe. Measured on a field case: a defect both spring-audit and impact-chain called medium composes to high (8.06) once reachable unauthenticated and writes to the database are in it. ask risk . · ask risk . --min-band high · ask risk . --limit 10reference

6 · Developer Workflows

The everyday loop: diff-based PR review, symptom-driven bug triage, and delta context for continuous agent runs. ask review-pr · ask fix-bug · ask prepare-contextreference

7 · Utilities

ask rename-class (word-boundary Java rename) · ask chunk-file (split large files for agents) · ask cache (status / warm / model / clear / freshness) → reference


Command tiers

A tier says what an output is worth relying on — it is a stability promise, not a value ranking, and not the pricing tier (Free/Pro gates repository size, never capability). posture is the most differentiated command in the product and it is experimental: both are true, and they are two different facts.

Tier Promise Commands
core contract stable within a major — safe to gate CI on endpoints · spring-audit · migrate-check · impact · impact-chain · pr-impact · verify
supported maintained; fields are added, never removed without a major every command not named in another row
experimental shape may change in a minor — do not gate CI on it risk · posture · archetype
parked kept working, no longer developed retrieve

The same table is printed by ask --help, and both are generated from one authority (cli.COMMAND_TIERS) — the battery fails if a command is in no tier, or if this file drifts from it.


Every command, in one table

ask --help shows a short header; this is the full surface. If you only read one row, read posture.

Command Tier Answers Note
posture experimental which beans a profile set wires, and how effective endpoint access differs between two sets the most differentiated capability here
risk experimental what each defect actually costs, once reach, access and write effect are in it defect_severity × reachability × auth_verdict × write_effect; every factor names its authority
endpoints core every REST endpoint, effective path, security policy, confidence Spring MVC + JAX-RS (~65 % recall on JAX-RS sub-resource locators)
spring-audit core transactional anomalies + security surface + validation gaps --ci, -f github-comment
migrate-check core Boot 2→3 readiness: located blockers, per-dimension score, effort --blast-radius orders the re-test plan
impact / impact-chain core blast radius of a change, to the endpoints it reaches target the interface, not the Impl
pr-impact core the same, scoped to a PR diff gating command: --fail-on, exit codes
verify core does the repo satisfy its declared contracts, relative to a baseline .ask/contracts.yml; exit 0/1/2
verify-edit supported did the working-tree edits change runtime behaviour semantic diff gate for the edit loop
--compact / --agent bounded structural context for an agent flags of the root command, not commands: not tiered
onboard / explain / cold-start supported orientation in an unfamiliar repo; per-class summary; bootstrap snapshot
export / repo-ir / schema supported tool-agnostic views (C4, module graph, integrations); symbol-level IR; published JSON Schemas
modernize supported coupling hubs, cycles, dead zones, refactor candidates
review-pr / fix-bug / prepare-context supported diff review, symptom triage, task-shaped context
plan / compare / delta / contract-diff supported what to review for a change; candidates by measured cost; outcome of a change; public-contract break no verdicts, measured cost only
validation supported request-body validation coverage and gaps
baseline capture|diff · trend supported versioned architectural metrics over time; ask trend <dir> reads the series trend reporting, not gating. Baselines land in .ask/baselines inside the repository — the history travels with the code, not with a vendor; an audit that must not write passes --dir. Automate it: baseline-ci.yml
retrieve parked typed knowledge queries over the model
archetype experimental evidence-based architectural archetype
rename-class / chunk-file supported word-boundary Java rename; split a large file for an agent
cache status|warm|model|clear · auth · telemetry · mcp · config · version supported housekeeping activate too

What it does — and doesn't

ASK Engine reduces exploration cost. It accelerates context acquisition and computes blast radius; it does not replace reading code — it reduces how often an agent needs to. All signals are static and deterministic (annotations, import graph, file structure) — no runtime analysis, no LLM guessing.

Honest limits worth knowing before you rely on it:

  • impact on an implementation class (OrderServiceImpl) returns 0 callers in Spring Boot — callers inject the interface. Always target the interface.
  • no_security_signal on an endpoint means no recognized method-level annotation, not "unsecured" — Spring Security filter chains and custom authorization annotations show as no_security_signal unless taught via config (below).
  • spring-audit / impact-chain are Java/Spring only; non-Java repos return spring_detected: false.
  • Event topology (--type events) resolves Spring ApplicationEvent / @EventListener chains only — not Kafka/RabbitMQ/Redis routes.
  • Architecture classification is tuned for Spring MVC layered apps; SPI/plugin models (e.g. Quarkus extensions) may be misclassified. JAX-RS subresource-locator endpoint recall is ~65%.
  • Self-invocation @Transactional bypass (same-class call skipping the proxy) is not detected.

What the security surface does not answer

Published because a product that states its limits is not compared on breadth — it is compared on depth. Each row is emitted in the payload too (non_coverage), so an agent reading JSON sees the same boundary a buyer reads here.

Not covered Why What answers it
Whether request input reaches a sink — no dataflow or taint analysis. This engine resolves structure and wiring without compiling. Taint needs value flow through a program, which is a different analysis with a different failure mode: an unsound one produces confident findings that are wrong, and every claim here is meant to be checkable against the line that produced it. A dataflow scanner (Semgrep, CodeQL). What this product adds on top is reachability: which of that tool's findings sit behind an endpoint that is reachable unauthenticated.
Secrets outside Java and Spring configuration — a credential in a deployment descriptor, a Helm value, a CI variable file. The file population this analyzer reads is the Java source and the Spring configuration convention. A secret elsewhere is not missed by a weak rule; it is outside the set of files anything here opens. A dedicated secret scanner over the whole tree (gitleaks, trufflehog).
Filter-chain order and per-filter URL patterns — the presence of a custom filter is structural only. Which filter runs first is decided by bean ordering this analyzer does not resolve. Where two active configurations both match a request, the answer published is undecided rather than a guess. ask posture --profile <set> states per endpoint what the readable rules decide and what they leave undecided.
Known vulnerabilities in dependencies — no CVE database, no version advisory matching. Enriching a vulnerability feed is a different product with a different update cadence; a stale embedded database is worse than no database, because it reads as a clean bill of health. A dependency scanner (Trivy, OWASP Dependency-Check). impact-chain then answers which of its findings anything actually reaches.

Two more boundaries worth stating in the same voice:

Not covered Why What answers it
Applying a migration — nothing here edits source. This is the diagnosis layer: it measures what must change and what each change would reach. Rewriting code is an execution problem with an established executor, and duplicating it would mean maintaining a second, worse one. OpenRewrite. migrate-check publishes the recipe each finding carries (recipes[]), which is the input that executor takes.
What the process environment sets at start-up — the profile set, properties and secrets a container is given. Nothing in a repository can observe the environment of a process that has not started. A value read from a file here is the default the repository ships, never a guarantee of what runs. ask posture --resolve-environments reads every artefact in the repository that names the profile set, says when none of them decides it, and ranks the admissible sets by what each leaves open.

Positioning. Until an executor ships, this is the diagnosis layer: it measures what must change, what each change reaches, and what a gate should block — and it removes none of it. Field evaluation scored it 7/10 as a report generator and 5.5/10 as a development tool, and that gap is the honest description, not a defect to argue with. migrate-check publishes the OpenRewrite recipe each finding carries so the executor that does apply changes has its input.


Pricing

🎉 Early-adoption: Pro is currently unlocked for everyone. Every install runs with full Pro entitlements — no size gate, no key. The tiers below describe the model the paywall will return to later.

What that means concretely. Ask the product: ask auth status answers in one block — entitlement (what runs today), source (why — a licence, the unlock, or the free tier), authenticated (whether a credential exists, which is a separate fact) and when_it_changes (what you lose when that source stops applying). A fresh install reads entitlement: pro, source: early_adoption_unlock, authenticated: false — unauthenticated and entitled, stated as two facts instead of one contradiction. When the unlock ends, gating returns by repo size and automation, never by command: posture, endpoints, spring-audit and migrate-check stay in the base tier at full output. Nothing you can run today becomes a paid-only command tomorrow.

Gating is by repo size and automation — never by command. Every command runs at full power on Free for small and mid-size repos; you upgrade when the work gets bigger or automated.

Free — €0 Pro — €19/mo · €190/yr per dev
Repo size ≤ 500 Java source files > 500 Java files (enterprise monoliths)
Commands All of them, full output Same commands, unlocked at scale
impact / fix-bug / review-pr / modernize ✅ full on small repos ✅ full on large repos (Free gets a capped preview)
prepare-context delta 30 free runs/repo unlimited — CI/CD automation
MCP local server, offline, no data egress

Non-Java repos are free at any size — the size limit counts Java source files only. ASK Engine monetises enterprise Java monoliths. Activate with ask activate <key>. Full breakdown: docs/PRODUCT_TIERS.md.


Configuration & privacy

ask config              # version, config file path, telemetry status
ask telemetry enable    # anonymous telemetry is OFF by default (opt-in)

Nothing is collected or transmitted unless you turn telemetry on — not on the first run, not in CI. If you do opt in, it collects version, OS, commands, flags, duration, repo-size range, and errors: no source code, paths, secrets, or output. Turn it off again with ask telemetry disable, export SOURCECODE_TELEMETRY=0, or DO_NOT_TRACK=1.

Auditing someone else's code — regulated, client-owned or public-sector? Nothing to do. Telemetry was opt-out until 3.3.0; it is opt-in now (docs/DEFECT-LEDGER.md P-1), because a default you must remember to disable is the wrong default for third-party code. An explicit choice you made before is unchanged.

Custom security annotations. Teach endpoints, spring-audit, and explain about project-specific authorization annotations via an optional sourcecode.config.json at the repo root (otherwise they report policy: "none_detected"):

{
  "customSecurityAnnotations": [
    { "fullyQualifiedName": "com.example.security.CustomSecurityAnnotation", "shortName": "CustomSecurityAnnotation" }
  ]
}

Matching endpoints report policy: "custom" and drop out of the no_security_signal count.


Documentation

Doc What it covers
USER_GUIDE.md Full command reference, flags, output schema, workflows
contracts.md Declare invariants in .ask/contracts.yml (or derive them with ask verify --init), enforce them in the edit loop and in CI
baseline-ci.yml Capture an architectural baseline per release, so the series exists when you want to read it
posture.md What a profile set actually wires — down to which endpoints its filter chain permits — and what could not be decided (experimental)
migrate-check.md Migration rule catalogue (MIG-001..043) + Hibernate stratification
MODERNIZATION.md The modernization product: assess → understand → plan → execute
PRODUCT_TIERS.md Free vs Pro, pricing model
DEMO-5MIN.md A reproducible 5-minute demo
MANUAL-USUARIO.md Guía de usuario en español
PRODUCT_IDENTITY.md ask (command) vs sourcecode (package/alias)
privacy.md Telemetry and data-handling policy
DEFECT-LEDGER.md Every defect found in the field, its class, and which release closed it — published on purpose

Project details


Release history Release notifications | RSS feed

This version

4.0.1

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sourcecode-4.0.1.tar.gz (1.2 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

sourcecode-4.0.1-py3-none-any.whl (1.2 MB view details)

Uploaded Python 3

File details

Details for the file sourcecode-4.0.1.tar.gz.

File metadata

  • Download URL: sourcecode-4.0.1.tar.gz
  • Upload date:
  • Size: 1.2 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for sourcecode-4.0.1.tar.gz
Algorithm Hash digest
SHA256 183454d3315dbbf9d785928a9c1be5af204da798c9376170cf0a1a4d1ab279ea
MD5 a6ed2eb27db87d95414b9770c43f1e87
BLAKE2b-256 b9f845477ca224c76f517a8fea316419280d1c82179daeba048c70790f8cff71

See more details on using hashes here.

File details

Details for the file sourcecode-4.0.1-py3-none-any.whl.

File metadata

  • Download URL: sourcecode-4.0.1-py3-none-any.whl
  • Upload date:
  • Size: 1.2 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for sourcecode-4.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 73f2be672862076f21f82c2d411ac3c82242e7d8624b68a6a54e3d48be291a1a
MD5 ca00c504d2b0ae85697ba47b90b23b13
BLAKE2b-256 08790ab05b87b9eb205ab8e01321d05015e395e256b114ab0d4667841495830b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page