Skip to main content

Count tokens in git-tracked files using tiktoken encoding for LLM API cost estimation

Project description

Count Tokens

A fast CLI tool that counts tokens in files using tiktoken encoding, automatically respecting .gitignore patterns. Perfect for estimating LLM API costs before processing your codebase.

Why Count Tokens?

When working with LLMs, understanding token usage helps you:

  • Estimate API costs before sending large codebases to models
  • Stay within context limits for different LLM models
  • Optimize prompts by identifying the largest files
  • Budget effectively for AI-assisted development workflows

Installation

From PyPI

pip install count-tokens-cli

From Source

git clone https://github.com/StGerman/count-tokens.git
cd count-tokens
poetry install

Quick Start

# Count tokens in your current repository
count-tokens

# Get just the summary without file breakdown
count-tokens --no-files

# Focus on specific file types
count-tokens --extensions .py .js .ts .md

Usage Examples

Basic Token Counting

count-tokens
Top 4 files by token count:
============================================================
   1,209 tokens | src/main.py
     639 tokens | docs/api.md
     290 tokens | README.md
     138 tokens | config.yaml
============================================================

Total tokens: 2,276
Total files: 4
Average per file: 569 tokens

Cost estimates (per API call):
  Claude Sonnet 4    $0.0068
  Claude Opus 4      $0.0341
  GPT-4o             $0.0057
  GPT-4o mini        $0.0003
  GPT-4 Turbo        $0.0228

Summary Only

count-tokens --no-files

Perfect for CI/CD pipelines or when you just need the totals.

Custom File Types

count-tokens --extensions .py .js .md

Focus on specific languages or documentation files.

Different Encodings

count-tokens --encoding o200k_base

Use different tiktoken encodings (cl100k_base, o200k_base, etc.).

Show More Files

count-tokens --top 50

Display up to 50 files instead of the default 30.

Supported File Types

The tool automatically detects and counts tokens in 25+ common file extensions:

Languages: .py .js .ts .jsx .tsx .java .go .rb .php .c .cpp .h .cs .swift .kt .rs .scala .r .m .sh

Config & Data: .yml .yaml .json .xml .toml .sql

Web & Docs: .html .css .scss .vue .md

Features

Gitignore Aware - Automatically respects .gitignore patterns
No Git Required - Works in any directory, not just git repositories
Smart Filtering - Supports 25+ code file extensions
Cost Estimation - Real-time pricing for popular LLM APIs
Flexible Encodings - Works with all tiktoken encodings
Top Files View - Identify your largest files quickly
Error Resilient - Gracefully handles unreadable files
Fast Performance - Optimized for large repositories

Development

Setup

# Install dependencies
poetry install

# Run during development
poetry run count-tokens
# or make executable and run directly
chmod +x count_tokens.py
./count_tokens.py

Building & Publishing

# Build package
poetry build

# Publish to PyPI
poetry publish

Testing

# Test on current repository
poetry run count-tokens

# Test with different options
poetry run count-tokens --no-files --encoding o200k_base

Requirements

  • Python 3.11+
  • tiktoken library for token encoding
  • pathspec library for .gitignore pattern matching

How It Works

The tool walks through all files in the current directory and subdirectories, automatically excluding files that match patterns in:

  • .gitignore files (current directory and parent directories)
  • Common ignore patterns (.git/ directories, etc.)

This means you get clean token counts without build artifacts, dependencies, or other files you typically don't want to include in LLM contexts.

Contributing

Contributions welcome! Please feel free to submit issues or pull requests.

License

MIT License - see LICENSE file for details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

count_tokens_cli-0.1.0.tar.gz (5.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

count_tokens_cli-0.1.0-py3-none-any.whl (6.3 kB view details)

Uploaded Python 3

File details

Details for the file count_tokens_cli-0.1.0.tar.gz.

File metadata

  • Download URL: count_tokens_cli-0.1.0.tar.gz
  • Upload date:
  • Size: 5.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/1.8.3 CPython/3.11.3 Darwin/25.0.0

File hashes

Hashes for count_tokens_cli-0.1.0.tar.gz
Algorithm Hash digest
SHA256 870ef404a3fcd4462daa268fb37927e945bed72ab3914006976e782d7c1a48db
MD5 555792864e90cb426613db4fb41221b5
BLAKE2b-256 aa460e421fd2433547de18ed888ff79df6a732de9abe66782da217cad9f18a40

See more details on using hashes here.

File details

Details for the file count_tokens_cli-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: count_tokens_cli-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 6.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/1.8.3 CPython/3.11.3 Darwin/25.0.0

File hashes

Hashes for count_tokens_cli-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 21c9a661da7272013ddaffbf8007e03ee59a2951349c065f19c0212fb81e3635
MD5 95a1d336b384c48c7e508d5c2e1a1d36
BLAKE2b-256 d22c43171132c27dde4285898ce742389450b0c961bb27dd117dbc5550875385

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page