Skip to main content

ai-benchmark-analyzer

PyPI version License: MIT Downloads LinkedIn

A Python package that analyzes and summarizes AI benchmark results from unstructured text descriptions. It extracts structured insights like top-performing models, key metrics, and comparative analysis using pattern-matching techniques.


📦 Installation

Install via pip:

pip install ai_benchmark_analyzer

🚀 Quick Start

Basic Usage (Default LLM7)

from ai_benchmark_analyzer import ai_benchmark_analyzer

user_input = """
Model X achieved 92.3% accuracy on SOTA dataset with 128B parameters.
Model Y scored 89.1% but used only 7B parameters.
"""
response = ai_benchmark_analyzer(user_input)
print(response)

Custom LLM Integration

Replace the default ChatLLM7 with your preferred LLM (e.g., OpenAI, Anthropic, Google):

OpenAI Example

from langchain_openai import ChatOpenAI
from ai_benchmark_analyzer import ai_benchmark_analyzer

llm = ChatOpenAI()
response = ai_benchmark_analyzer(user_input, llm=llm)

Anthropic Example

from langchain_anthropic import ChatAnthropic
from ai_benchmark_analyzer import ai_benchmark_analyzer

llm = ChatAnthropic()
response = ai_benchmark_analyzer(user_input, llm=llm)

Google Generative AI Example

from langchain_google_genai import ChatGoogleGenerativeAI
from ai_benchmark_analyzer import ai_benchmark_analyzer

llm = ChatGoogleGenerativeAI()
response = ai_benchmark_analyzer(user_input, llm=llm)

🔧 Parameters

Parameter Type Description
user_input str Raw text describing benchmark results (required).
api_key Optional[str] LLM7 API key (auto-fetched from LLM7_API_KEY env var if not provided).
llm Optional[BaseChatModel] Custom LLM instance (defaults to ChatLLM7).

🔗 Default LLM: LLM7

The package uses ChatLLM7 by default. Free tier rate limits are sufficient for most use cases. For higher limits:

  • Set LLM7_API_KEY environment variable.
  • Pass the key directly: ai_benchmark_analyzer(api_key="your_key").

Get a free API key at LLM7 Token.


📝 Output Format

The function returns a list of structured strings matching a predefined regex pattern, ensuring consistent and reliable output formatting. Example output:

[
    "Model: ModelX, Metric: Accuracy, Value: 92.3%, Dataset: SOTA, Parameters: 128B",
    "Model: ModelY, Metric: Accuracy, Value: 89.1%, Dataset: SOTA, Parameters: 7B"
]

🛠️ Customization

  • Pattern Matching: The output adheres to a regex pattern (defined in prompts.py). Modify this file to adjust expected output formats.
  • LLM Prompts: System/human prompts are configurable via prompts.py.

📜 License

MIT


📢 Support & Issues

For bugs/feature requests, open an issue on GitHub.


👤 Author

Eugene Evstafev (@chigwell) 📧 hi@euegne.plus

Metadata

Release files for ai-benchmark-analyzer 2025.12.21193050

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ai-benchmark-analyzer 2025.12.21193050
File Size Uploaded
ai_benchmark_analyzer-2025.12.21193050.tar.gz 4.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ai-benchmark-analyzer 2025.12.21193050
File Interpreter ABI Platform
ai_benchmark_analyzer-2025.12.21193050-py3-none-any.whl Python 3 none any Details

Total release size: 10.1 kB

Release files / ai_benchmark_analyzer-2025.12.21193050.tar.gz

Download URL ai_benchmark_analyzer-2025.12.21193050.tar.gz
Size 4.7 kB
Tags Source
SHA-256 checksum
How to use checksums
e560cacc9dc7205cd1ee18ad1a3c28ed1f610b6baee6d175776ebd7a04bb7b82
BLAKE2b-256 checksum
How to use checksums
6fe9729e879974ef32c4ca427b3750dc845c2898b3515797285a20e5cae97680
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.1

Release files / ai_benchmark_analyzer-2025.12.21193050-py3-none-any.whl

Download URL ai_benchmark_analyzer-2025.12.21193050-py3-none-any.whl
Size 5.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e4547c58f1ee46fecc9939e502298ab7593a05d053098b33779b09053d77250a
BLAKE2b-256 checksum
How to use checksums
e812036613e2b65a194d8f27bd098f2fa080f213403612b4604949c60b55b709
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.1

Release history Release notifications | RSS feed

This release

2025.12.21193050 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page