Skip to main content

Advanced LLM token analysis and statistics toolkit for various data formats

Project description

Deus LLM Token Stats Guru

Advanced LLM token analysis and statistics toolkit for various data formats.

Python 3.10+ License: MIT

Features

  • Accurate Token Counting: Uses OpenAI's tiktoken library for precise token counts
  • Recursive Processing: Automatically finds and processes all CSV files in directories
  • Multiple Encoding Models: Supports different OpenAI models (gpt-4, gpt-3.5-turbo, etc.)
  • Dual CLI Interface: Two convenient command-line tools
  • Comprehensive Output: Detailed statistics including file sizes, row counts, and processing times
  • JSON Export: Results can be exported to JSON format
  • Type Safety: Full type hints and modern Python features
  • Extensible: Designed for future support of multiple data formats beyond CSV

Installation

pip install deus-llm-token-stats-guru

Quick Start

Command Line Usage

# Count tokens in current directory
deus-llm-token-guru .

# Alternative command
llm-token-stats ./data

# Count tokens using specific model
deus-llm-token-guru /path/to/csv/files --model gpt-3.5-turbo

# Save results to JSON file
llm-token-stats /path/to/csv/files --output results.json

# Enable debug logging
deus-llm-token-guru /path/to/csv/files --debug --log-file debug.log

# Quiet mode (suppress progress output)
llm-token-stats ./data --quiet

Python API Usage

from deus_llm_token_stats_guru import CSVTokenCounter
from pathlib import Path

# Initialize counter
counter = CSVTokenCounter(encoding_model="gpt-4")

# Count tokens in a single CSV file
result = counter.count_tokens_in_csv(Path("data.csv"))
print(f"Total tokens: {result['total_tokens']:,}")

# Process entire directory
summary = counter.count_tokens_in_directory(Path("./csv_files"))
print(f"Processed {summary['total_files']} files")
print(f"Total tokens: {summary['total_tokens']:,}")

Available Commands

The package provides two CLI commands:

  • deus-llm-token-guru: Main command name
  • llm-token-stats: Alternative shorter command

Both commands have identical functionality.

API Reference

CSVTokenCounter

Main class for counting tokens in CSV files.

Methods

  • __init__(encoding_model: str = "gpt-4"): Initialize with specific encoding model
  • count_tokens_in_text(text: str) -> int: Count tokens in a text string
  • count_tokens_in_csv(file_path: Path) -> CountResult: Count tokens in single CSV file
  • count_tokens_in_directory(directory: Path) -> CountSummary: Process all CSV files in directory

Type Definitions

class CountResult(TypedDict):
    file_path: str
    total_tokens: int
    row_count: int
    column_count: int
    encoding_model: str
    file_size_bytes: int

class CountSummary(TypedDict):
    total_files: int
    total_tokens: int
    total_rows: int
    total_file_size_bytes: int
    encoding_model: str
    file_results: List[CountResult]
    processing_time_seconds: float

Examples

Basic File Processing

from deus_llm_token_stats_guru import CSVTokenCounter
import pandas as pd
from pathlib import Path

# Create sample data
data = {
    'title': ['AI Introduction', 'Machine Learning'],
    'content': [
        'Artificial intelligence is transforming industries',
        'ML enables computers to learn from data'
    ]
}

# Save to CSV
df = pd.DataFrame(data)
df.to_csv('articles.csv', index=False)

# Count tokens
counter = CSVTokenCounter()
result = counter.count_tokens_in_csv(Path('articles.csv'))

print(f"File: articles.csv")
print(f"Tokens: {result['total_tokens']:,}")
print(f"Rows: {result['row_count']}")
print(f"Columns: {result['column_count']}")

Directory Processing with Different Models

from deus_llm_token_stats_guru import CSVTokenCounter
from pathlib import Path

models = ["gpt-4", "gpt-3.5-turbo"]
directory = Path("./data")

for model in models:
    counter = CSVTokenCounter(encoding_model=model)
    summary = counter.count_tokens_in_directory(directory)

    print(f"Model: {model}")
    print(f"Total tokens: {summary['total_tokens']:,}")
    print(f"Processing time: {summary['processing_time_seconds']:.2f}s")
    print()

Cost Estimation

from deus_llm_token_stats_guru import CSVTokenCounter

counter = CSVTokenCounter(encoding_model="gpt-4")
summary = counter.count_tokens_in_directory("./training_data")

# Estimate OpenAI API costs (example rates)
gpt4_cost_per_1k = 0.03  # USD per 1K tokens
total_cost = (summary['total_tokens'] / 1000) * gpt4_cost_per_1k
print(f"Estimated GPT-4 cost: ${total_cost:.2f}")

CLI Output Format

The CLI tools output JSON with the following structure:

{
  "summary": {
    "total_files": 3,
    "total_tokens": 15420,
    "total_rows": 150,
    "total_file_size_mb": 2.1,
    "encoding_model": "gpt-4",
    "processing_time_seconds": 0.85
  },
  "file_details": [
    {
      "file_path": "/path/to/file1.csv",
      "tokens": 5140,
      "rows": 50,
      "columns": 4,
      "size_mb": 0.7
    }
  ]
}

Environment Setup

1. Create Virtual Environment

python -m venv .venv

2. Activate Virtual Environment

# Linux/macOS
source .venv/bin/activate

# Windows
.venv\Scripts\activate

3. Install Package

pip install --upgrade pip
pip install deus-llm-token-stats-guru

Development

Local Installation

git clone https://github.com/yourusername/deus-llm-token-stats-guru.git
cd deus-llm-token-stats-guru
pip install -e .

Running Tests

# Run all tests
pytest

# Run with coverage
pytest --cov=src --cov-report=term-missing

# Run specific test file
pytest tests/unit/test_core.py

Code Quality

# Format code
ruff format src/ tests/

# Lint code
ruff check src/ tests/

# Type checking
mypy src/

Building Package

# Build package
python setup.py sdist bdist_wheel

# Check package
twine check dist/*

# Test installation
pip install dist/*.whl

Supported Models

The package supports all OpenAI tiktoken encoding models:

  • gpt-4 (default)
  • gpt-3.5-turbo
  • text-davinci-003
  • text-davinci-002
  • code-davinci-002
  • Custom encodings via tiktoken

Performance

  • Processes ~1000 rows/second on typical hardware
  • Memory usage scales with CSV file size
  • Supports files with millions of rows
  • Recursive directory scanning with progress tracking

Error Handling

The package includes comprehensive error handling:

  • FileProcessingError: Issues reading or processing CSV files
  • EncodingError: Problems with tiktoken encoding
  • ConfigurationError: Invalid configuration or paths

Use Cases

Training Data Analysis

# Analyze ML training datasets
deus-llm-token-guru ./training_data --model gpt-4 --output analysis.json

Cost Planning

# Estimate API costs for large datasets
llm-token-stats ./production_data --model gpt-3.5-turbo --output cost_analysis.json

Batch Processing

# Process multiple directories
from deus_llm_token_stats_guru import CSVTokenCounter

counter = CSVTokenCounter()
directories = ["./data1", "./data2", "./data3"]

for dir_path in directories:
    summary = counter.count_tokens_in_directory(dir_path)
    print(f"{dir_path}: {summary['total_tokens']:,} tokens")

Future Roadmap

  • Support for additional data formats (JSON, XML, TXT)
  • Integration with more LLM providers
  • Advanced analytics and reporting features
  • Web interface for easy visualization

Contributing

  1. Fork the repository
  2. Create a feature branch: git checkout -b feature-name
  3. Make changes with tests
  4. Run quality checks: ruff check && mypy src/ && pytest
  5. Submit a pull request

License

This project is licensed under the MIT License - see the LICENSE file for details.

Changelog

v0.1.0

  • Initial release
  • Core CSV token counting functionality
  • Dual CLI interface (deus-llm-token-guru and llm-token-stats)
  • Support for multiple encoding models
  • Comprehensive test suite
  • JSON output format
  • Python API with full type hints

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

deus-llm-token-stats-guru-0.2.0.tar.gz (21.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

deus_llm_token_stats_guru-0.2.0-py3-none-any.whl (25.2 kB view details)

Uploaded Python 3

File details

Details for the file deus-llm-token-stats-guru-0.2.0.tar.gz.

File metadata

File hashes

Hashes for deus-llm-token-stats-guru-0.2.0.tar.gz
Algorithm Hash digest
SHA256 576f610ae315d643ee83776fd2fbf6c7f3a53d2cc90b887234186f89b3520cfd
MD5 b338ffe73e73c1b153b41ccffcc9b864
BLAKE2b-256 344e342188945d3d0a4fbf1bad7b009ded06093cef85ae604379ed2c67215a4c

See more details on using hashes here.

File details

Details for the file deus_llm_token_stats_guru-0.2.0-py3-none-any.whl.

File metadata

File hashes

Hashes for deus_llm_token_stats_guru-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 513f406003f6024324d60b55b041d276c992ee1aa8c7a65e328ea46046dc8b8e
MD5 e144d3a0ff1dbf579b8229979ad3e395
BLAKE2b-256 861ff7652fb4cddd630adb0d8bce3a89775d14e3825c301221e4e148dc17f3c3

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page