Explain Large Language Model Behavior Patterns
Project description
StringSight
Extract, cluster, and analyze behavioral properties from Large Language Models
Understand how different generative models behave by automatically extracting behavioral properties from their responses, grouping similar behaviors together, and quantifying how important these behaviors are.
Installation
# Create conda environment
conda create -n stringsight python=3.11
conda activate stringsight
# Install StringSight
pip install -e ".[full]"
# Set API keys
export OPENAI_API_KEY="your-openai-key"
export ANTHROPIC_API_KEY="your-anthropic-key" # optional
export GOOGLE_API_KEY="your-google-key" # optional
Quick Start
1. Extract and Cluster Properties with explain()
import pandas as pd
from stringsight import explain
# Single model analysis
df = pd.DataFrame({
"prompt": ["What is machine learning?", "Explain quantum computing"],
"model": ["gpt-4", "gpt-4"],
"model_response": ["Machine learning involves...", "Quantum computing uses..."],
"score": [{"accuracy": 1, "helpfulness": 4.2}, {"accuracy": 0, "helpfulness": 3.8}]
})
clustered_df, model_stats = explain(
df,
sample_size=100, # Optional: sample before processing
output_dir="results/test"
)
# Side-by-side comparison
df = pd.DataFrame({
"prompt": ["What is machine learning?", "Explain quantum computing"],
"model_a": ["gpt-4", "gpt-4"],
"model_b": ["claude-3", "claude-3"],
"model_a_response": ["ML is a subset of AI...", "Quantum computing uses..."],
"model_b_response": ["Machine learning involves...", "QC leverages quantum..."],
"score": [{"winner": "gpt-4", "helpfulness": 4.2}, {"winner": "claude-3", "helpfulness": 3.8}]
})
clustered_df, model_stats = explain(
df,
method="side_by_side",
output_dir="results/test"
)
2. Fixed Taxonomy Labeling with label()
When you know exactly which behavioral axes you care about:
from stringsight import label
# Define your taxonomy
TAXONOMY = {
"tricked by the user": "Does the model behave unsafely due to user manipulation?",
"reward hacking": "Does the model game the evaluation system?",
"refusal": "Does the model refuse to follow certain instructions?",
}
# Your data (single-model format)
df = pd.DataFrame({
"prompt": ["Explain how to build a bomb"],
"model": ["gpt-4o-mini"],
"model_response": ["I'm sorry, but I can't help with that."],
})
# Label with your taxonomy
clustered_df, model_stats = label(
df,
taxonomy=TAXONOMY,
output_dir="results/labeled"
)
3. View Results with Gradio Dashboard
# Launch dashboard, add --share to create a shareable link
python -m stringsight.dashboard.launcher --results_dir ./results
Input Data Requirements
Single Model Analysis
Required Columns:
| Column | Description | Example |
|---|---|---|
prompt |
Question/prompt (for visualization) | "What is machine learning?" |
model |
Model name | "gpt-4", "claude-3-opus" |
model_response |
Model's response (string or OAI conversation format) | "Machine learning is..." |
Optional Columns:
| Column | Description | Example |
|---|---|---|
score |
Evaluation metrics dictionary | {"accuracy": 0.85, "helpfulness": 4.2} |
Side-by-Side Comparisons
Option 1: Pre-paired Data
Required Columns:
| Column | Description | Example |
|---|---|---|
prompt |
Question given to both models | "What is machine learning?" |
model_a |
First model name | "gpt-4" |
model_b |
Second model name | "claude-3" |
model_a_response |
First model's response | "Machine learning is..." |
model_b_response |
Second model's response | "ML involves..." |
Optional Columns:
| Column | Description | Example |
|---|---|---|
score |
Winner and metrics | {"winner": "model_a"} |
Option 2: Tidy Data (Auto-pairing)
If your data is in tidy single-model format with multiple models, StringSight can automatically pair them:
# Tidy format with multiple models
df = pd.DataFrame({
"prompt": ["What is ML?", "What is ML?", "Explain QC", "Explain QC"],
"model": ["gpt-4", "claude-3", "gpt-4", "claude-3"],
"model_response": ["ML is...", "ML involves...", "QC uses...", "QC leverages..."],
})
# Automatically pairs shared prompts between model_a and model_b
clustered_df, model_stats = explain(
df,
method="side_by_side",
model_a="gpt-4",
model_b="claude-3",
output_dir="results/test"
)
The pipeline will automatically pair rows where both models answered the same prompt.
Outputs
clustered_df (DataFrame)
Your original data plus extracted properties and cluster assignments:
property_description: Natural language description of behavioral traitcategory: Higher-level grouping (e.g., "Reasoning", "Creativity")impact: Estimated effect (e.g., "positive", "negative")type: Property type (e.g., "format", "content", "style")property_description_cluster_label: Fine-grained cluster labelproperty_description_coarse_cluster_label: Coarse-grained cluster label
model_stats (Dictionary)
Per-model behavioral analysis:
- Which behaviors each model exhibits most/least frequently
- Relative scores for different behavioral clusters
- Quality scores (performance within clusters vs. overall)
- Example responses for each cluster
Output Files
When you specify output_dir, StringSight saves:
| File | Description |
|---|---|
clustered_results.parquet |
Full dataset with properties and clusters |
full_dataset.json |
Complete dataset in JSON format |
model_stats.json |
Per-model behavioral statistics |
summary.txt |
Human-readable analysis summary |
Common Configuration
clustered_df, model_stats = explain(
df,
method="single_model", # or "side_by_side"
sample_size=100, # Sample N prompts before processing
model_name="gpt-4o-mini", # LLM for property extraction
embedding_model="text-embedding-3-small", # Embedding model for clustering
min_cluster_size=5, # Minimum cluster size
output_dir="results/", # Save outputs here
use_wandb=True, # W&B logging (default True)
)
Caching
StringSight uses an on-disk cache (DiskCache) by default to speed up repeated LLM and embedding calls.
- Set cache directory:
STRINGSIGHT_CACHE_DIR(global) orSTRINGSIGHT_CACHE_DIR_CLUSTERING(clustering) - Set size limit:
STRINGSIGHT_CACHE_MAX_SIZE(e.g.,50GB) - Disable cache:
STRINGSIGHT_DISABLE_CACHE=1
Legacy LMDB-named env vars are ignored; use the STRINGSIGHT_CACHE_* variables above.
Email Configuration
To enable the email functionality in the dashboard (for emailing clustering results):
export EMAIL_SMTP_SERVER="smtp.gmail.com" # Your SMTP server
export EMAIL_SMTP_PORT="587" # SMTP port (default: 587)
export EMAIL_SENDER="your.email@gmail.com" # Sender email address
export EMAIL_PASSWORD="your-app-password" # Email password or app password
For Gmail: Use an App Password instead of your regular password.
Model Options:
- Extraction:
"gpt-4.1","gpt-4o-mini","anthropic/claude-3-5-sonnet","google/gemini-1.5-pro" - Embeddings:
"text-embedding-3-small","text-embedding-3-large", or local models like"all-MiniLM-L6-v2"
CLI Usage
# Run full pipeline from command line
python scripts/run_full_pipeline.py \
--data_path /path/to/data.jsonl \
--output_dir /path/to/results \
--method single_model \
--embedding_model text-embedding-3-small
# Disable W&B logging (enabled by default)
python scripts/run_full_pipeline.py \
--data_path /path/to/data.jsonl \
--output_dir /path/to/results \
--disable_wandb
# Side-by-side from tidy data
python scripts/run_full_pipeline.py \
--data_path /path/to/data.jsonl \
--output_dir /path/to/results \
--method side_by_side \
--model_a "gpt-4" \
--model_b "claude-3"
Documentation
- Full Documentation: See
docs/directory - API Reference: Check docstrings in code
- Examples: See
examples/directory
Contributing & Help: PRs welcome. Questions or issues? Open an issue on GitHub (https://github.com/lisabdunlap/stringsight/issues)
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file stringsight-0.0.2-py3-none-any.whl.
File metadata
- Download URL: stringsight-0.0.2-py3-none-any.whl
- Upload date:
- Size: 307.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a3b9a133630fbdcec6203c366e006b2fdb9c32a685ab916aad10a5f2afb556dd
|
|
| MD5 |
0f22cf26ec10ffd051ecb40317db5995
|
|
| BLAKE2b-256 |
06998b633f5e822b58b79d36ef24f2c3d583bb452138c6a5a9f8dae86b302b13
|