Skip to main content

A tool for mining sentences from games. Update: Dependencies, replay buffer based line searching, and bug fixes.

Project description

gamesentenceminer

GSM (GameSentenceMiner)

Turn your gaming time into language mastery.


🎮 See it in Action

Demo Gif

  • OCR to get get text from a game that doesn't support text hooks.
  • Look up words with Yomitan in game.
  • Create Anki cards with game audio + screenshot (or gif) automatically.

What does it do?

GSM is an application designed to automate the process of creating flashcards while you play. It sits between your game and Anki, handling audio recording, screenshots, and OCR so you don't have to interrupt your gameplay.

📝 Anki Card Enhancement

GSM automatically adds context to your Anki cards whenever you create them.

  • Audio Capture: Uses Voice Activity Detection (VAD) to record and trim the specific voice line associated with the text.
  • Screenshots: Captures the game state the moment the line is spoken. GIFs and Black Bar Removal are supported.
  • Mine from History: Go back and create cards from previous lines you've encountered (i.e. cutscenes).
  • Multi-Line Support: Capture multiple lines of dialogue into one card using the built-in Texthooker.
  • AI Translation: Optional integration to provide sentence translations using your own API key.

https://github.com/user-attachments/assets/df6bc38e-d74d-423e-b270-8a82eec2394c

👁️ OCR (Text Recognition)

For games that don't have a text hook (Agent/Textractor), GSM uses a custom fork of OwOCR to read text directly from the screen.

This opens up all kinds of posssibilities for games that would otherwise be inaccessible for language learning/sentence mining. For example I've made cards with games like Metal Gear Solid 1+2, Titanfall 2, and Sekiro, all using GSM's OCR.

  • Easy Setup: Managed installation means you don't need to fiddle with terminals.
  • Two-Pass System: Clean, fast output similar to as if you had a hook.
  • Customizable Capture Zones: Define exactly where text appears on your screen for optimal results.

https://github.com/user-attachments/assets/07240472-831a-40e6-be22-c64b880b0d66

🖥️ Overlay

GSM includes a transparent overlay for instant dictionary lookups.

  • Hover over characters in-game to see definitions via Yomitan.
  • Create cards without ever leaving the game window.
  • Automatically Generated Furigana Display In Game.

Overlay Demo

📊 Statistics

Track your immersion habits with the stats dashboard.

  • Kanji Grid: View every Kanji you've encountered and click them to see their source sentences.
  • Goals: Set daily reading targets.
  • Tools: Clean up and organize your mining history.

stats


🚀 Getting Started

  1. Download: Get the latest release.
  2. Install: Watch the Installation Guide.
  3. Requirements:
    • An Anki tool (Yomitan, JL, etc.)
    • A text source (Agent, Textractor, or GSM's built-in OCR)
    • A game

📚 Documentation

For full setup guides and configuration details, check the Wiki (Currently WIP).

❤️ Acknowledgements

Integrated Components

This project includes modified versions of the following libraries, I got tired of submodule hell so I've included them directly here for easier management all credits go to the original authors:

Star History

Star History Chart

Sponsors

Free code signing provided by SignPath.io, certificate by SignPath Foundation.

Project details


Release history Release notifications | RSS feed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gamesentenceminer-2026.6.11.tar.gz (28.9 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

gamesentenceminer-2026.6.11-py3-none-any.whl (29.1 MB view details)

Uploaded Python 3

File details

Details for the file gamesentenceminer-2026.6.11.tar.gz.

File metadata

  • Download URL: gamesentenceminer-2026.6.11.tar.gz
  • Upload date:
  • Size: 28.9 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.15

File hashes

Hashes for gamesentenceminer-2026.6.11.tar.gz
Algorithm Hash digest
SHA256 40bb118c4342cb2c5618cb4457a0c9fdd73919e22828f24b89e3eb27d08f96ac
MD5 e8085b4d144f30a1babb769224ee6dd5
BLAKE2b-256 ea3d3a24aa968bdf6de2eb2e4e46d4a5fa23007cdcb7b5f0872a3d026d0bc0b8

See more details on using hashes here.

File details

Details for the file gamesentenceminer-2026.6.11-py3-none-any.whl.

File metadata

File hashes

Hashes for gamesentenceminer-2026.6.11-py3-none-any.whl
Algorithm Hash digest
SHA256 1d98744fa320d9db30585e86bff3154fb87fb8bea969eba3c2a7a92975b92d23
MD5 b0e45612b5c7e63d3565deb191eda551
BLAKE2b-256 27d8c0ffe1ede58aa5f661f3afb59e1d5b112faf83beffa992fc17b49ea079fc

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page