topicscout
Find the GitHub repositories on a topic, and next time only the new ones.
pip install topicscout # or run it without installing: uvx topicscout run github-scrapers
topicscout run github-scrapers
19 found, 17 kept, 17 new -> scout-github-scrapers/new.md
new.md is a table of what appeared since the last run, most stars first: stars,
last push, language, license, a guess at the interface from the README (MCP, pip,
npm, REST, Docker...) and the description. candidates.json keeps everything seen,
so a weekly run is a short list, not the same 300 repos again.
No dependencies (the MCP server is an optional extra), Python 3.11+. GITHUB_TOKEN is optional: without it the searches
are paced to stay under GitHub's 10 per minute and READMEs are capped at 40 per run.
Why
GitHub search is good at "repos with this topic" and bad at "what's new since I last looked, minus the forks, the awesome-lists, the abandoned ones and the repos that only tagged themselves". topicscout runs a fixed set of searches, applies your filters, remembers what it has seen and what you already have.
It was the discovery step of adebench, a benchmark for AI memory systems, which uses it to find new memories to measure.
Your own topic
Copy a built-in profile (they are in topicscout/profiles/),
change the queries and the words a hit must contain, and run it:
topicscout run my-topic.json
Run it again next week: new.md lists only what appeared in between. Put the repos
you already know or rejected in scout-my-topic/known.txt.
From a script or an agent
--json prints the result on stdout, progress stays on stderr:
topicscout run github-scrapers --json -q
{"profile": "github-scrapers", "found": 19, "kept": 17, "new": [{"full_name": "yusufkaraaslan/Skill_Seekers", "stars": 15052, "interface": ["cli", "pip", "action"], "...": "..."}, "..."], "stopped": null, "out": "scout-github-scrapers"}
It is also a Python library (from topicscout.core import Profile, run) and an MCP server.
As an MCP server
pip install "topicscout[mcp]"
claude mcp add topicscout -- topicscout mcp # Claude Code; any MCP client runs `topicscout mcp`
Or in a client's JSON config: {"command": "topicscout", "args": ["mcp"], "env": {"GITHUB_TOKEN": "..."}}.
| tool | what it does | network |
|---|---|---|
list_profiles |
built-in profiles and the state folder | no |
scout_run |
search GitHub with a profile, return what is new | yes, slow without a token |
scout_candidates |
what past runs saw: filter by status (new, seen, known), text, stars |
no |
scout_mark_known |
add a repo to known.txt with a note, so it never comes back as new |
no |
doctor |
token and remaining GitHub quota | yes |
The state of each profile lives in ~/.topicscout/<profile> (or $TOPICSCOUT_HOME);
every tool also takes out, the same folder the CLI's --out takes, so the agent and
the command line can share one state.
Commands
topicscout run PROFILE [--out DIR] [--min-stars N] [--days N] [--readme-max N]
[--known-from GLOB ...] [--json] [-q]
topicscout profiles # built-in profiles
topicscout doctor # token and remaining GitHub quota
topicscout mcp # MCP server on stdio (needs topicscout[mcp])
PROFILEis a built-in name or a path to your own JSON profile.--outis the state folder (defaultscout-<profile>):candidates.json,new.md, and yourknown.txt.--jsonprints the summary and the new repos as JSON on stdout; progress and errors always go to stderr, so scripts and agents can read stdout as is.
What counts as known
Repos you already have are marked known and never listed as new:
known.txtin the state folder: oneowner/repoper line,#comments. Put rejected repos here too.--known-from "adapters/*.py": the GitHub URLs in the first 20 lines of those files. adebench points it at its adapters, so a memory with an adapter drops out.
Profiles
Built in: ai-memory (memory systems for AI agents and LLMs) and github-scrapers
(tools that scrape, crawl or mine GitHub). A profile is a JSON file:
{
"name": "ai-memory",
"description": "Memory systems for AI agents and LLMs",
"queries": ["topic:agent-memory", "\"llm memory\" in:name,description"],
"exclude": "(?i)awesome|curated list|\\bpapers\\b",
"require": [
{"pattern": "(?i)memor|\\bmem\\b", "own_words": true},
{"pattern": "(?i)\\b(agents?|llms?|ai|mcp)\\b"}
],
"interfaces": [["mcp", "(?i)\\bmcp\\b"], ["pip", "(?i)\\bpip install\\b"]],
"min_stars": 10,
"days": 90
}
queries: GitHub repository search queries. Each one always getsfork:false archived:false stars:>=min_stars pushed:>=today-days; up to 300 results per query.exclude: a regex on name, description and topics; a match drops the repo.require: every pattern must match. Withown_wordsit must be in the name or description, not only in a topic: a database that tagged itselfagent-memoryis not a memory. Hyphens and underscores count as spaces.interfaces:[label, regex]pairs tried on the README of each new repo. Leave it out and no README is fetched.require_language(default true) drops repos with no language, which are usually lists and docs.
Patterns are Python regexes; add (?i) for case-insensitive.
Rate limits
On a 403 or 429 it waits for the reset if that is under 70 seconds; otherwise it
stops, saves what it found and says so in new.md and on stderr. The next run
picks up the READMEs it did not get to.
How it compares
- ghcrawl goes deep into one repository (issues and PRs, embeddings, clusters); topicscout goes wide across GitHub for a topic.
- top-github-scraper lists the top repos for a keyword, once; topicscout adds filters, profiles and memory across runs.
License
MIT
Metadata
Release files for topicscout 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| topicscout-0.2.0.tar.gz | 20.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| topicscout-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 36.8 kB
Release files / topicscout-0.2.0.tar.gz
| Download URL | topicscout-0.2.0.tar.gz |
|---|---|
| Size | 20.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
429d81437a420f4c4fe6363030c34f6921ba4e28352a490f1b3f14f7349f003c
|
|
BLAKE2b-256 checksum How to use checksums |
898f52fd83b01f384d99bc3c1fac74af3b4c0ef0d3b38fbdf3e5023f0d5e3731
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / topicscout-0.2.0-py3-none-any.whl
| Download URL | topicscout-0.2.0-py3-none-any.whl |
|---|---|
| Size | 16.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a32465ffb5ca9f77a83cb86fcffe1d37dd2676af647ae6161cb4e4d72025c2da
|
|
BLAKE2b-256 checksum How to use checksums |
8a0221099461afa8abc7f4d0a12f43a2f573f263ad86fd285d707a6fd665995c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log