Skip to main content

Persistent structural context and ultra-fast repeated analysis for AI coding agents

Project description

ASK Engine

ASK — Actionable Software Knowledge. Persistent structural intelligence for AI coding agents.

Context · Impact · Migration · Architecture · Review — everything from one structural model.

Version Python

ASK Engine is the product. The CLI command is ask. The legacy sourcecode command still works as a deprecated alias (it prints a one-line notice and forwards to ask) and remains the Python/PyPI package name for now. The authoritative version is whatever ask version reports. See docs/PRODUCT_IDENTITY.md.


The problem

Every time an AI coding agent starts a new session, it has to re-parse the repository from scratch. For a large Java or TypeScript monolith, that means 5–15 seconds per invocation. Multiply by dozens of agent turns per hour, and repo context acquisition becomes a real bottleneck — not just latency, but tokens, compute, and iteration velocity.

ASK Engine solves this with a persistent structural cache keyed on file content hashes. After the first scan, every subsequent invocation returns pre-built context in milliseconds. The repo doesn't change? The cache doesn't expire.

The cache is not a performance optimization. It is what makes ASK Engine usable as infrastructure rather than a one-off tool.


Proof — measured on real repos

Repo Size Cold scan Cache hit Speedup
Keycloak 7,885 Java files 10.5s 0.6s ~17x
BroadleafCommerce 2,985 Java files 2.7s 0.3s ~9x

Cache keyed on content hashes — invalidated only when source changes. On repeated agent sessions against the same codebase, nearly every invocation is a cache hit.

At 0.3s per call, ASK Engine becomes constant infrastructure inside agent loops — call it before every edit, every PR review, every test run, without batching or caching manually.

What a warm actually covers. ask cache warm runs the compact analysis: it rebuilds the shared structural layers (L1/L2 + the Repository Intelligence Snapshot + the shared Canonical IR) and the compact view. Pass --agent to warm the agent view as well. Deeper projections are separate keys and are not covered by either — --full, --env-map and a raised --depth recompute on first use, as do most prepare-context tasks. Measured on a 3,342-file Spring monolith: cache warm 103s → --compact --git-context 1s (hit), but --agent --full --env-map --depth 20 still 171s (miss).

One rule for invalidation, and a table that names every command. Every layer keys on the exact tree state: any change to the analysed files invalidates it, committed or not. Which of your commands a warm helps — the answer, only the shared work, or nothing — is published per command: ask cache model, or docs/CACHE.md.


Install

# Homebrew (macOS / Linux)
brew tap haroundominique/sourcecode && brew install sourcecode

# pip / pipx
pipx install sourcecode        # or: pip install sourcecode

ask version                    # ask 2.5.1

Package vs. command. The install package is named sourcecode this release (renaming the distribution is a separate, breaking change). Installing it gives you the canonical ask command plus the deprecated sourcecode alias.


Quickstart

Start with these four. In three independent field evaluations they carried most of the measured value, and posture --diff is the one no evaluator found an equivalent for — commercial or open source.

# What does each profile set ACTUALLY wire — and what changes between them?
# Resolves conditional beans and the filter chain, then diffs effective endpoint access.
ask posture /path/to/repo --diff dev:prod

# Every REST endpoint with its effective path (context-path + servlet path resolved),
# its inferred security policy and a confidence per endpoint.
ask endpoints /path/to/repo

# Spring semantic audit: transactional anomalies (private @Transactional = silent
# CGLIB no-op), security surface, request-body validation.
ask spring-audit /path/to/repo

# Spring Boot 2→3 readiness: located blockers, per-dimension score, effort estimate.
ask migrate-check /path/to/repo --compact

Then the everyday loop:

# High-signal structural summary — warm cache ~0.3s, cold 2–10s
ask --compact

# Blast radius: what breaks if this class changes?  (target the INTERFACE, not the Impl)
ask impact OrderService /path/to/repo

# Onboard to an unfamiliar codebase
ask onboard /path/to/repo

# PR review: risk, test gaps, changed modules
ask review-pr /path/to/repo --since main

# CI gate on NEW violations only, instead of on pre-existing debt
ask verify /path/to/repo --capture-baseline   # accept today's debt, once
ask verify /path/to/repo                      # then: only new violations block

Adopting a gate on a real codebase. A repository that starts declaring contracts already violates them somewhere; ask verify is baseline-relative by default (--fail-on new) so the gate survives contact with reality instead of being switched off on day one. ask baseline capture|diff|trend is a different thing: versioned architectural metrics over time, for trend reporting rather than blocking.

Full command reference: docs/USER_GUIDE.md · posture in depth: docs/posture.md.


Capabilities

Everything is computed from one cached structural model. Seven groups:

1 · Structural Context

Bounded, noise-free repo context designed to drop straight into an agent's context window. ask --compact · ask --agent · ask onboard · ask cold-startreference

2 · Impact Analysis

Blast radius from a class or interface — reverse dependencies, through Spring DI, to the HTTP endpoints a change reaches. ask impact · ask impact-chain (TX/SEC-enriched) · ask pr-impactreference

3 · Architecture Intelligence

The system map: module graph, dependency views, REST surface, per-class summaries, and a symbol-level IR for downstream tooling. ask export · ask repo-ir · ask endpoints · ask explainreference

4 · Migration & Modernization

Is this codebase ready to upgrade? Per-dimension readiness (Jakarta / Spring Boot / JDK / Hibernate), located blockers, an effort estimate, and — with --blast-radius — the endpoints whose call path runs through each blocker, so the re-test plan is ordered by regression scope. ask migrate-check · ask modernizemigrate-check reference · MODERNIZATION.md

5 · Spring Analysis

Deterministic Spring semantics: transactional anomalies (e.g. @Transactional on a private method = silent CGLIB no-op), security surface, request-body validation. ask spring-audit · ask validationreference

5b · Runtime Posture (experimental — and the most differentiated thing here)

What a profile set actually wires: which conditional beans register, which do not, and which conditions could not be decided at all — then the effective endpoint access that follows from the filter chain. --diff answers the question nobody else answers in one command: what changes between dev and prod, across every endpoint at once. ask posture · ask posture --diff dev:prod · ask posture --property k=vposture.md

Unresolved is a first-class outcome: a condition the resolver cannot decide is reported as a hole with the condition named, never folded into active or inactive. A posture answer that guesses is a confident security falsehood — the worst failure mode this tool has.

6 · Developer Workflows

The everyday loop: diff-based PR review, symptom-driven bug triage, and delta context for continuous agent runs. ask review-pr · ask fix-bug · ask prepare-contextreference

7 · Utilities

ask rename-class (word-boundary Java rename) · ask chunk-file (split large files for agents) · ask cache (status / warm / model / clear / freshness) → reference


Command tiers

A tier says what an output is worth relying on — it is a stability promise, not a value ranking, and not the pricing tier (Free/Pro gates repository size, never capability). posture is the most differentiated command in the product and it is experimental: both are true, and they are two different facts.

Tier Promise Commands
core contract stable within a major — safe to gate CI on endpoints · spring-audit · migrate-check · impact · impact-chain · pr-impact · verify
supported maintained; fields are added, never removed without a major every command not named in another row
experimental shape may change in a minor — do not gate CI on it posture · archetype
parked kept working, no longer developed retrieve

The same table is printed by ask --help, and both are generated from one authority (cli.COMMAND_TIERS) — the battery fails if a command is in no tier, or if this file drifts from it.


Every command, in one table

ask --help shows a short header; this is the full surface. If you only read one row, read posture.

Command Tier Answers Note
posture experimental which beans a profile set wires, and how effective endpoint access differs between two sets the most differentiated capability here
endpoints core every REST endpoint, effective path, security policy, confidence Spring MVC + JAX-RS (~65 % recall on JAX-RS sub-resource locators)
spring-audit core transactional anomalies + security surface + validation gaps --ci, -f github-comment
migrate-check core Boot 2→3 readiness: located blockers, per-dimension score, effort --blast-radius orders the re-test plan
impact / impact-chain core blast radius of a change, to the endpoints it reaches target the interface, not the Impl
pr-impact core the same, scoped to a PR diff gating command: --fail-on, exit codes
verify core does the repo satisfy its declared contracts, relative to a baseline .ask/contracts.yml; exit 0/1/2
verify-edit supported did the working-tree edits change runtime behaviour semantic diff gate for the edit loop
--compact / --agent bounded structural context for an agent flags of the root command, not commands: not tiered
onboard / explain / cold-start supported orientation in an unfamiliar repo; per-class summary; bootstrap snapshot
export / repo-ir / schema supported tool-agnostic views (C4, module graph, integrations); symbol-level IR; published JSON Schemas
modernize supported coupling hubs, cycles, dead zones, refactor candidates
review-pr / fix-bug / prepare-context supported diff review, symptom triage, task-shaped context
plan / compare / delta / contract-diff supported what to review for a change; candidates by measured cost; outcome of a change; public-contract break no verdicts, measured cost only
validation supported request-body validation coverage and gaps
baseline capture|diff · trend supported versioned architectural metrics over time; ask trend <dir> reads the series trend reporting, not gating. Baselines land in .ask/baselines inside the repository — the history travels with the code, not with a vendor; an audit that must not write passes --dir. Automate it: baseline-ci.yml
retrieve parked typed knowledge queries over the model
archetype experimental evidence-based architectural archetype
rename-class / chunk-file supported word-boundary Java rename; split a large file for an agent
cache status|warm|model|clear · auth · telemetry · mcp · config · version supported housekeeping activate too

What it does — and doesn't

ASK Engine reduces exploration cost. It accelerates context acquisition and computes blast radius; it does not replace reading code — it reduces how often an agent needs to. All signals are static and deterministic (annotations, import graph, file structure) — no runtime analysis, no LLM guessing.

Honest limits worth knowing before you rely on it:

  • impact on an implementation class (OrderServiceImpl) returns 0 callers in Spring Boot — callers inject the interface. Always target the interface.
  • no_security_signal on an endpoint means no recognized method-level annotation, not "unsecured" — Spring Security filter chains and custom authorization annotations show as no_security_signal unless taught via config (below).
  • spring-audit / impact-chain are Java/Spring only; non-Java repos return spring_detected: false.
  • Event topology (--type events) resolves Spring ApplicationEvent / @EventListener chains only — not Kafka/RabbitMQ/Redis routes.
  • Architecture classification is tuned for Spring MVC layered apps; SPI/plugin models (e.g. Quarkus extensions) may be misclassified. JAX-RS subresource-locator endpoint recall is ~65%.
  • Self-invocation @Transactional bypass (same-class call skipping the proxy) is not detected.

What the security surface does not answer

Published because a product that states its limits is not compared on breadth — it is compared on depth. Each row is emitted in the payload too (non_coverage), so an agent reading JSON sees the same boundary a buyer reads here.

Not covered Why What answers it
Whether request input reaches a sink — no dataflow or taint analysis. This engine resolves structure and wiring without compiling. Taint needs value flow through a program, which is a different analysis with a different failure mode: an unsound one produces confident findings that are wrong, and every claim here is meant to be checkable against the line that produced it. A dataflow scanner (Semgrep, CodeQL). What this product adds on top is reachability: which of that tool's findings sit behind an endpoint that is reachable unauthenticated.
Secrets outside Java and Spring configuration — a credential in a deployment descriptor, a Helm value, a CI variable file. The file population this analyzer reads is the Java source and the Spring configuration convention. A secret elsewhere is not missed by a weak rule; it is outside the set of files anything here opens. A dedicated secret scanner over the whole tree (gitleaks, trufflehog).
Filter-chain order and per-filter URL patterns — the presence of a custom filter is structural only. Which filter runs first is decided by bean ordering this analyzer does not resolve. Where two active configurations both match a request, the answer published is undecided rather than a guess. ask posture --profile <set> states per endpoint what the readable rules decide and what they leave undecided.
Known vulnerabilities in dependencies — no CVE database, no version advisory matching. Enriching a vulnerability feed is a different product with a different update cadence; a stale embedded database is worse than no database, because it reads as a clean bill of health. A dependency scanner (Trivy, OWASP Dependency-Check). impact-chain then answers which of its findings anything actually reaches.

Two more boundaries worth stating in the same voice:

Not covered Why What answers it
Applying a migration — nothing here edits source. This is the diagnosis layer: it measures what must change and what each change would reach. Rewriting code is an execution problem with an established executor, and duplicating it would mean maintaining a second, worse one. OpenRewrite. migrate-check publishes the recipe each finding carries (recipes[]), which is the input that executor takes.
What the process environment sets at start-up — the profile set, properties and secrets a container is given. Nothing in a repository can observe the environment of a process that has not started. A value read from a file here is the default the repository ships, never a guarantee of what runs. ask posture --resolve-environments reads every artefact in the repository that names the profile set, says when none of them decides it, and ranks the admissible sets by what each leaves open.

Positioning. Until an executor ships, this is the diagnosis layer: it measures what must change, what each change reaches, and what a gate should block — and it removes none of it. Field evaluation scored it 7/10 as a report generator and 5.5/10 as a development tool, and that gap is the honest description, not a defect to argue with. migrate-check publishes the OpenRewrite recipe each finding carries so the executor that does apply changes has its input.


Pricing

🎉 Early-adoption: Pro is currently unlocked for everyone. Every install runs with full Pro entitlements — no size gate, no key. The tiers below describe the model the paywall will return to later.

What that means concretely. Ask the product: ask auth status answers in one block — entitlement (what runs today), source (why — a licence, the unlock, or the free tier), authenticated (whether a credential exists, which is a separate fact) and when_it_changes (what you lose when that source stops applying). A fresh install reads entitlement: pro, source: early_adoption_unlock, authenticated: false — unauthenticated and entitled, stated as two facts instead of one contradiction. When the unlock ends, gating returns by repo size and automation, never by command: posture, endpoints, spring-audit and migrate-check stay in the base tier at full output. Nothing you can run today becomes a paid-only command tomorrow.

Gating is by repo size and automation — never by command. Every command runs at full power on Free for small and mid-size repos; you upgrade when the work gets bigger or automated.

Free — €0 Pro — €19/mo · €190/yr per dev
Repo size ≤ 500 Java source files > 500 Java files (enterprise monoliths)
Commands All of them, full output Same commands, unlocked at scale
impact / fix-bug / review-pr / modernize ✅ full on small repos ✅ full on large repos (Free gets a capped preview)
prepare-context delta 30 free runs/repo unlimited — CI/CD automation
MCP local server, offline, no data egress

Non-Java repos are free at any size — the size limit counts Java source files only. ASK Engine monetises enterprise Java monoliths. Activate with ask activate <key>. Full breakdown: docs/PRODUCT_TIERS.md.


Configuration & privacy

ask config              # version, config file path, telemetry status
ask telemetry enable    # anonymous telemetry is OFF by default (opt-in)

Nothing is collected or transmitted unless you turn telemetry on — not on the first run, not in CI. If you do opt in, it collects version, OS, commands, flags, duration, repo-size range, and errors: no source code, paths, secrets, or output. Turn it off again with ask telemetry disable, export SOURCECODE_TELEMETRY=0, or DO_NOT_TRACK=1.

Auditing someone else's code — regulated, client-owned or public-sector? Nothing to do. Telemetry was opt-out until 3.3.0; it is opt-in now (docs/DEFECT-LEDGER.md P-1), because a default you must remember to disable is the wrong default for third-party code. An explicit choice you made before is unchanged.

Custom security annotations. Teach endpoints, spring-audit, and explain about project-specific authorization annotations via an optional sourcecode.config.json at the repo root (otherwise they report policy: "none_detected"):

{
  "customSecurityAnnotations": [
    { "fullyQualifiedName": "com.example.security.CustomSecurityAnnotation", "shortName": "CustomSecurityAnnotation" }
  ]
}

Matching endpoints report policy: "custom" and drop out of the no_security_signal count.


Documentation

Doc What it covers
USER_GUIDE.md Full command reference, flags, output schema, workflows
contracts.md Declare invariants in .ask/contracts.yml (or derive them with ask verify --init), enforce them in the edit loop and in CI
baseline-ci.yml Capture an architectural baseline per release, so the series exists when you want to read it
posture.md What a profile set actually wires — down to which endpoints its filter chain permits — and what could not be decided (experimental)
migrate-check.md Migration rule catalogue (MIG-001..043) + Hibernate stratification
MODERNIZATION.md The modernization product: assess → understand → plan → execute
PRODUCT_TIERS.md Free vs Pro, pricing model
DEMO-5MIN.md A reproducible 5-minute demo
MANUAL-USUARIO.md Guía de usuario en español
PRODUCT_IDENTITY.md ask (command) vs sourcecode (package/alias)
privacy.md Telemetry and data-handling policy
DEFECT-LEDGER.md Every defect found in the field, its class, and which release closed it — published on purpose

Project details


Release history Release notifications | RSS feed

This version

3.8.0

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sourcecode-3.8.0.tar.gz (1.2 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

sourcecode-3.8.0-py3-none-any.whl (1.2 MB view details)

Uploaded Python 3

File details

Details for the file sourcecode-3.8.0.tar.gz.

File metadata

  • Download URL: sourcecode-3.8.0.tar.gz
  • Upload date:
  • Size: 1.2 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for sourcecode-3.8.0.tar.gz
Algorithm Hash digest
SHA256 8c29ab8c27d0c31f8e9ec69176cbab061c42bc8912bc8540efaa477c9e7046a0
MD5 a849eea34c48e7253c5e8883cb91727f
BLAKE2b-256 1e2ac2b79753f16ed924ef1cfe25b40e01092565b7697dc7a60357e399d599c0

See more details on using hashes here.

File details

Details for the file sourcecode-3.8.0-py3-none-any.whl.

File metadata

  • Download URL: sourcecode-3.8.0-py3-none-any.whl
  • Upload date:
  • Size: 1.2 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.6

File hashes

Hashes for sourcecode-3.8.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e9f4048e1e0465e90121740f1f5d298e455db5c4f38a9074d6b94cdb163b2d1a
MD5 bfa7dae9e77464059921946a4f443a3c
BLAKE2b-256 ba6aa77199feb8a6a9f32dd0af7873a6278e4033197f3ff782180470e9e2b075

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page