Skip to main content

topicscout

ci PyPI

Find the GitHub repositories on a topic, and next time only the new ones.

pip install topicscout        # or run it without installing: uvx topicscout run github-scrapers
topicscout run github-scrapers
19 found, 17 kept, 17 new -> scout-github-scrapers/new.md

new.md is a table of what appeared since the last run, most stars first: stars, last push, language, license, a guess at the interface from the README (MCP, pip, npm, REST, Docker...) and the description. candidates.json keeps everything seen, so a weekly run is a short list, not the same 300 repos again.

No dependencies, Python 3.11+. GITHUB_TOKEN is optional: without it the searches are paced to stay under GitHub's 10 per minute and READMEs are capped at 40 per run.

Why

GitHub search is good at "repos with this topic" and bad at "what's new since I last looked, minus the forks, the awesome-lists, the abandoned ones and the repos that only tagged themselves". topicscout runs a fixed set of searches, applies your filters, remembers what it has seen and what you already have.

It was the discovery step of adebench, a benchmark for AI memory systems, which uses it to find new memories to measure.

Your own topic

Copy a built-in profile (they are in topicscout/profiles/), change the queries and the words a hit must contain, and run it:

topicscout run my-topic.json

Run it again next week: new.md lists only what appeared in between. Put the repos you already know or rejected in scout-my-topic/known.txt.

From a script or an agent

--json prints the result on stdout, progress stays on stderr:

topicscout run github-scrapers --json -q
{"profile": "github-scrapers", "found": 19, "kept": 17, "new": [{"full_name": "yusufkaraaslan/Skill_Seekers", "stars": 15052, "interface": ["cli", "pip", "action"], "...": "..."}, "..."], "stopped": null, "out": "scout-github-scrapers"}

It is a command-line tool and a Python library (from topicscout.core import Profile, run), not an MCP server; an agent calls it like any other command.

Commands

topicscout run PROFILE [--out DIR] [--min-stars N] [--days N] [--readme-max N]
                       [--known-from GLOB ...] [--json] [-q]
topicscout profiles    # built-in profiles
topicscout doctor      # token and remaining GitHub quota
  • PROFILE is a built-in name or a path to your own JSON profile.
  • --out is the state folder (default scout-<profile>): candidates.json, new.md, and your known.txt.
  • --json prints the summary and the new repos as JSON on stdout; progress and errors always go to stderr, so scripts and agents can read stdout as is.

What counts as known

Repos you already have are marked known and never listed as new:

  • known.txt in the state folder: one owner/repo per line, # comments. Put rejected repos here too.
  • --known-from "adapters/*.py": the GitHub URLs in the first 20 lines of those files. adebench points it at its adapters, so a memory with an adapter drops out.

Profiles

Built in: ai-memory (memory systems for AI agents and LLMs) and github-scrapers (tools that scrape, crawl or mine GitHub). A profile is a JSON file:

{
  "name": "ai-memory",
  "description": "Memory systems for AI agents and LLMs",
  "queries": ["topic:agent-memory", "\"llm memory\" in:name,description"],
  "exclude": "(?i)awesome|curated list|\\bpapers\\b",
  "require": [
    {"pattern": "(?i)memor|\\bmem\\b", "own_words": true},
    {"pattern": "(?i)\\b(agents?|llms?|ai|mcp)\\b"}
  ],
  "interfaces": [["mcp", "(?i)\\bmcp\\b"], ["pip", "(?i)\\bpip install\\b"]],
  "min_stars": 10,
  "days": 90
}
  • queries: GitHub repository search queries. Each one always gets fork:false archived:false stars:>=min_stars pushed:>=today-days; up to 300 results per query.
  • exclude: a regex on name, description and topics; a match drops the repo.
  • require: every pattern must match. With own_words it must be in the name or description, not only in a topic: a database that tagged itself agent-memory is not a memory. Hyphens and underscores count as spaces.
  • interfaces: [label, regex] pairs tried on the README of each new repo. Leave it out and no README is fetched.
  • require_language (default true) drops repos with no language, which are usually lists and docs.

Patterns are Python regexes; add (?i) for case-insensitive.

Rate limits

On a 403 or 429 it waits for the reset if that is under 70 seconds; otherwise it stops, saves what it found and says so in new.md and on stderr. The next run picks up the READMEs it did not get to.

How it compares

  • ghcrawl goes deep into one repository (issues and PRs, embeddings, clusters); topicscout goes wide across GitHub for a topic.
  • top-github-scraper lists the top repos for a keyword, once; topicscout adds filters, profiles and memory across runs.

License

MIT

Metadata

Release files for topicscout 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for topicscout 0.1.0
File Size Uploaded
topicscout-0.1.0.tar.gz 15.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for topicscout 0.1.0
File Interpreter ABI Platform
topicscout-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 29.0 kB

Release files / topicscout-0.1.0.tar.gz

Download URL topicscout-0.1.0.tar.gz
Size 15.9 kB
Tags Source
SHA-256 checksum
How to use checksums
5fc85fb4632d9ad011e784241b6f5b28a8bd8b304b2c8ebee96e18c506c82b8c
BLAKE2b-256 checksum
How to use checksums
a6d2359b7a22cdd89a4952e494c0b217dc8456d63336f1c197b9a4d7e8d1e59e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / topicscout-0.1.0-py3-none-any.whl

Download URL topicscout-0.1.0-py3-none-any.whl
Size 13.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f1ff6fe2f793579e79f16f306cdf1b09dd8c112e418b391956cdf532e3beecef
BLAKE2b-256 checksum
How to use checksums
2cfb9f9b2b880ff3f3fb0f7b2476ac72624e73965f3940717238ed010e12a7a3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page