Skip to main content

A tool for obfuscating text by manipulating token IDs while preserving token count and structure

Project description

LLM Token Obfuscator

A tool for obfuscating text by manipulating token IDs while preserving token count and structure. Originally developed for benchmarking LLM inference performance and prefix caching behavior by generating test data that maintains token patterns but with obscured text.

Overview

This project provides a system for obfuscating text by applying a shift to token IDs. The obfuscation is reversible and preserves the token count, making it useful for:

  • Testing LLM systems with obfuscated content
  • Benchmarking tokenization performance
  • Creating privacy-preserving datasets
  • Generating synthetic text with realistic token distributions

Example of obfuscated text patterns:

Original: The quick brown fox jumps
Obfuscated: eng($_ ét rl manga

Original: The quick brown fox runs
Obfuscated: eng($_ ét rl Android

Note how the common prefix "The quick brown" is obfuscated to "eng($_ ét" in both cases, preserving the pattern.

Installation

From PyPI (Recommended)

pip install llm-obfuscator

From Source

  1. Clone the repository:
git clone https://github.com/yourusername/llm-obfusicator.git
cd llm-obfusicator
  1. Install dependencies:
pip install -r requirements.txt
  1. Install the package in development mode:
pip install -e .

Usage

Command Line Interface

The package provides a command-line interface for easy use:

# Tokenize text
llm-obfuscator tokenize gpt-4 "Hello, world!"

# Obfuscate text
llm-obfuscator obfuscate gpt-4 "Hello, world!"

# Obfuscate text with a fixed shift
llm-obfuscator obfuscate gpt-4 "Hello, world!" --shift 42

Python API

from llm_obfuscator import obfuscate_text, tokenize_text

# Obfuscate text using a specific model's tokenizer
obfuscated = obfuscate_text("gpt-4", "Hello, world!")
print(obfuscated)

# Use a fixed shift value for deterministic results
obfuscated = obfuscate_text("gpt-4", "Hello, world!", shift=42)
print(obfuscated)

# Tokenize text
tokens = tokenize_text("gpt-4", "Hello, world!")
print(tokens)

Supported Models

The system supports both OpenAI and HuggingFace tokenizers:

  • OpenAI models: gpt-4, gpt-3.5-turbo, cl100k_base, etc.
  • HuggingFace models: gpt2, bert-base-uncased, etc.

Testing

The project includes several test suites to validate the obfuscation system:

Running All Tests

The easiest way to run all tests is to use the provided shell script:

# Make the script executable (if needed)
chmod +x run_all_tests.sh

# Run all tests
./run_all_tests.sh

This will run all test files in sequence, including unit tests and specialized test scripts.

Running Basic Tests

# Run all tests
python -m pytest tests/

# Run specific test file
python -m pytest tests/test.py

Specialized Test Scripts

The project includes specialized test scripts for different aspects of the obfuscation system:

# Test with real-world examples
python tests/test_real_world.py

# Test obfuscation demonstration
python tests/test_obfuscation.py

# Test mathematical properties
python tests/test_mapping_properties.py

Validation

The obfuscation system has been validated to ensure:

  1. Token Count Preservation: The number of tokens remains the same after obfuscation
  2. One-to-One Mapping: The obfuscation is a bijective (one-to-one) mapping
  3. Frequency Preservation: Token frequency distributions are preserved
  4. Reversibility: The original text can be recovered by applying the reverse shift

How It Works

The obfuscation process works as follows:

  1. Text is tokenized using the specified model's tokenizer
  2. Each token ID is shifted by a fixed amount (either specified or randomly generated)
  3. The shifted tokens are detokenized back to text

The shift operation is performed modulo the vocabulary size (typically 50,000) to ensure all tokens remain within the valid vocabulary range.

Note: The token count preservation has a known margin of error of up to 8% for obscured texts. We are working on improving this.

License

MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llm_obfuscator-0.1.0.tar.gz (19.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

llm_obfuscator-0.1.0-py3-none-any.whl (12.5 kB view details)

Uploaded Python 3

File details

Details for the file llm_obfuscator-0.1.0.tar.gz.

File metadata

  • Download URL: llm_obfuscator-0.1.0.tar.gz
  • Upload date:
  • Size: 19.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.0.1 CPython/3.9.19

File hashes

Hashes for llm_obfuscator-0.1.0.tar.gz
Algorithm Hash digest
SHA256 1b0f686a8ad7cbfa2953dea271d19051650b5decc9c513476fe14754646aed6b
MD5 01a3fc4c6bd2f6f9a0b6c8b99d906910
BLAKE2b-256 547ac313223f274e81c9fef963b7bae9fe7023cde07fe3352d455303e70b72e3

See more details on using hashes here.

File details

Details for the file llm_obfuscator-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: llm_obfuscator-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 12.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.0.1 CPython/3.9.19

File hashes

Hashes for llm_obfuscator-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e836e95ead3c03c3bda9b6813796c9ba0b85442a6b8cc6e0914be0d3bb851e35
MD5 e68b5183932fcf4f7f0032b874162820
BLAKE2b-256 f83ab835a40fe38184c79f8aa1be7e96177e1e27430c967454e3a8a1c7d96619

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page