Skip to main content

topicscout

ci PyPI

Find the GitHub repositories on a topic, and next time only the new ones.

pip install topicscout        # or run it without installing: uvx topicscout run github-scrapers
topicscout run github-scrapers
19 found, 17 kept, 17 new -> scout-github-scrapers/new.md

new.md is a table of what appeared since the last run, most stars first: stars, last push, language, license, a guess at the interface from the README (MCP, pip, npm, REST, Docker...) and the description. candidates.json keeps everything seen, so a weekly run is a short list, not the same 300 repos again.

No dependencies (the MCP server is an optional extra), Python 3.11+. GITHUB_TOKEN is optional: without it the searches are paced to stay under GitHub's 10 per minute and READMEs are capped at 40 per run.

Why

GitHub search is good at "repos with this topic" and bad at "what's new since I last looked, minus the forks, the awesome-lists, the abandoned ones and the repos that only tagged themselves". topicscout runs a fixed set of searches, applies your filters, remembers what it has seen and what you already have.

It was the discovery step of adebench, a benchmark for AI memory systems, which uses it to find new memories to measure.

Your own topic

Copy a built-in profile (they are in topicscout/profiles/), change the queries and the words a hit must contain, and run it:

topicscout run my-topic.json

Run it again next week: new.md lists only what appeared in between. Put the repos you already know or rejected in scout-my-topic/known.txt.

From a script or an agent

--json prints the result on stdout, progress stays on stderr:

topicscout run github-scrapers --json -q
{"profile": "github-scrapers", "found": 19, "kept": 17, "new": [{"full_name": "yusufkaraaslan/Skill_Seekers", "stars": 15052, "interface": ["cli", "pip", "action"], "...": "..."}, "..."], "stopped": null, "out": "scout-github-scrapers"}

It is also a Python library (from topicscout.core import Profile, run) and an MCP server.

As an MCP server

pip install "topicscout[mcp]"
claude mcp add topicscout -- topicscout mcp      # Claude Code; any MCP client runs `topicscout mcp`

Or in a client's JSON config: {"command": "topicscout", "args": ["mcp"], "env": {"GITHUB_TOKEN": "..."}}.

tool what it does network
list_profiles built-in profiles and the state folder no
scout_run search GitHub with a profile, return what is new yes, slow without a token
scout_candidates what past runs saw: filter by status (new, seen, known), text, stars no
scout_mark_known add a repo to known.txt with a note, so it never comes back as new no
doctor token and remaining GitHub quota yes

The state of each profile lives in ~/.topicscout/<profile> (or $TOPICSCOUT_HOME); every tool also takes out, the same folder the CLI's --out takes, so the agent and the command line can share one state.

Commands

topicscout run PROFILE [--out DIR] [--min-stars N] [--days N] [--readme-max N]
                       [--known-from GLOB ...] [--json] [-q]
topicscout profiles    # built-in profiles
topicscout doctor      # token and remaining GitHub quota
topicscout mcp         # MCP server on stdio (needs topicscout[mcp])
  • PROFILE is a built-in name or a path to your own JSON profile.
  • --out is the state folder (default scout-<profile>): candidates.json, new.md, and your known.txt.
  • --json prints the summary and the new repos as JSON on stdout; progress and errors always go to stderr, so scripts and agents can read stdout as is.

What counts as known

Repos you already have are marked known and never listed as new:

  • known.txt in the state folder: one owner/repo per line, # comments. Put rejected repos here too.
  • --known-from "adapters/*.py": the GitHub URLs in the first 20 lines of those files. adebench points it at its adapters, so a memory with an adapter drops out.

Profiles

Built in: ai-memory (memory systems for AI agents and LLMs) and github-scrapers (tools that scrape, crawl or mine GitHub). A profile is a JSON file:

{
  "name": "ai-memory",
  "description": "Memory systems for AI agents and LLMs",
  "queries": ["topic:agent-memory", "\"llm memory\" in:name,description"],
  "exclude": "(?i)awesome|curated list|\\bpapers\\b",
  "require": [
    {"pattern": "(?i)memor|\\bmem\\b", "own_words": true},
    {"pattern": "(?i)\\b(agents?|llms?|ai|mcp)\\b"}
  ],
  "interfaces": [["mcp", "(?i)\\bmcp\\b"], ["pip", "(?i)\\bpip install\\b"]],
  "min_stars": 10,
  "days": 90
}
  • queries: GitHub repository search queries. Each one always gets fork:false archived:false stars:>=min_stars pushed:>=today-days; up to 300 results per query.
  • exclude: a regex on name, description and topics; a match drops the repo.
  • require: every pattern must match. With own_words it must be in the name or description, not only in a topic: a database that tagged itself agent-memory is not a memory. Hyphens and underscores count as spaces.
  • interfaces: [label, regex] pairs tried on the README of each new repo. Leave it out and no README is fetched.
  • require_language (default true) drops repos with no language, which are usually lists and docs.

Patterns are Python regexes; add (?i) for case-insensitive.

Rate limits

On a 403 or 429 it waits for the reset if that is under 70 seconds; otherwise it stops, saves what it found and says so in new.md and on stderr. The next run picks up the READMEs it did not get to.

How it compares

  • ghcrawl goes deep into one repository (issues and PRs, embeddings, clusters); topicscout goes wide across GitHub for a topic.
  • top-github-scraper lists the top repos for a keyword, once; topicscout adds filters, profiles and memory across runs.

License

MIT

Metadata

Release files for topicscout 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for topicscout 0.2.0
File Size Uploaded
topicscout-0.2.0.tar.gz 20.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for topicscout 0.2.0
File Interpreter ABI Platform
topicscout-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 36.8 kB

Release files / topicscout-0.2.0.tar.gz

Download URL topicscout-0.2.0.tar.gz
Size 20.3 kB
Tags Source
SHA-256 checksum
How to use checksums
429d81437a420f4c4fe6363030c34f6921ba4e28352a490f1b3f14f7349f003c
BLAKE2b-256 checksum
How to use checksums
898f52fd83b01f384d99bc3c1fac74af3b4c0ef0d3b38fbdf3e5023f0d5e3731
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / topicscout-0.2.0-py3-none-any.whl

Download URL topicscout-0.2.0-py3-none-any.whl
Size 16.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a32465ffb5ca9f77a83cb86fcffe1d37dd2676af647ae6161cb4e4d72025c2da
BLAKE2b-256 checksum
How to use checksums
8a0221099461afa8abc7f4d0a12f43a2f573f263ad86fd285d707a6fd665995c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page