LLMDataLens
LLMDataLens is a powerful and flexible framework for evaluating LLM-based applications with structured output. It provides a comprehensive suite of tools for assessing the performance of language models across various metrics, with a focus on experiment tracking and reproducibility.
🌟 Features
- Structured Output Evaluation: Assess LLM outputs against ground truth data with precision.
- Customizable Metrics: Easily define and use custom metrics for comprehensive performance assessment.
- Experiment Tracking: Built-in experiment management for reproducibility and comparison.
- Prompt Versioning: Keep track of prompt evolution and its impact on model performance.
- Model Version Tracking: Monitor performance across different model versions.
- Flexible Integration: Seamlessly integrate with existing LLM pipelines and workflows.
- Extensible Architecture: Add custom metrics, evaluators, and experiment trackers with ease.
🚀 Quick Start
Installation
Install LLMDataLens directly from PyPI:
pip install llm-data-lens
For development or to get the latest version from the repository:
-
Clone the repository:
git clone https://github.com/codingmindset/LLMDataLens.git cd llmdatalens
-
Install the package using Poetry:
poetry install
Basic Usage
Here's a simple example to get you started:
from llmdatalens.evaluators import StructuredOutputEvaluator
from llmdatalens.core import LLMOutputData, GroundTruthData
from llmdatalens.core.metrics_registry import MetricNames
# Create an evaluator with specific metrics
evaluator = StructuredOutputEvaluator(
metrics=[MetricNames.OverallAccuracy, MetricNames.AverageLatency],
experiment_name="Invoice Processing Experiment"
)
# Add LLM output and ground truth data
llm_output = LLMOutputData(
raw_output="Processed invoice: $100",
structured_output={"invoice_amount": 100},
metadata={
"model_info": {"name": "GPT-3.5", "version": "1.0"},
"prompt_info": {"text": "Extract invoice amount:"}
}
)
ground_truth = GroundTruthData(
data={"invoice_amount": 100}
)
evaluator.add_llm_output(llm_output, latency=0.5, confidence=0.9)
evaluator.add_ground_truth(ground_truth)
# Evaluate
result = evaluator.evaluate()
# Print results
print(result.metrics)
# Access experiment data
experiment = evaluator.experiment_manager.get_experiment(evaluator.experiment_id)
print(f"Experiment: {experiment.name}")
print(f"Number of runs: {len(experiment.runs)}")
print(f"Prompts used: {len(experiment.prompts)}")
print(f"Models used: {list(experiment.models.keys())}")
📊 Advanced Features
Custom Metrics
Create and register custom metrics easily:
from llmdatalens.core.metrics_registry import register_metric
from llmdatalens.core.enums import MetricField
@register_metric("CustomF1Score", field=MetricField.Accuracy, input_keys=["y_true", "y_pred"])
def calculate_custom_f1_score(y_true, y_pred):
""" This description will be shown in the metrics registry """
# (Your custom F1 score calculation here
pass
Experiment Tracking
Track experiments, prompts, and model versions:
# Get prompt history
prompt_history = evaluator.experiment_manager.get_prompt_history(evaluator.experiment_id)
# Get model history
model_history = evaluator.experiment_manager.get_model_history(evaluator.experiment_id)
# Compare runs
for run in experiment.runs:
print(f"Run {run.id}: {run.metrics}")
For more detailed examples, check the examples/ directory in the repository. (More examples will be added soon!)
📘 Documentation
(Comming soon!)
🛠️ Project Structure
llmdatalens/
├── src/
│ └── llmdatalens/
│ ├── core/
│ │ ├── base_model.py
│ │ ├── enums.py
│ │ └── metrics_registry.py
│ ├── evaluators/
│ │ └── structured_output_evaluator.py
│ └── experiment/
│ ├── experiment_manager.py
│ └── models.py
├── tests/
│ ├── test_core/
│ ├── test_evaluators/
│ └── test_experiment/
├── examples/
├── docs/
├── pyproject.toml
└── README.md
🤝 Contributing
We welcome contributions to LLMDataLens! Here's how you can help:
- Fork the repository
- Create a new branch (
git checkout -b feature/AmazingFeature) - Make your changes
- Commit your changes (
git commit -m 'Add some AmazingFeature') - Push to the branch (
git push origin feature/AmazingFeature) - Open a Pull Request
Please read our Contributing Guidelines for more details.
📄 License
LLMDataLens is released under the MIT License. See the LICENSE file for details.
📬 Contact
If you have any questions, suggestions, or just want to say hi, feel free to reach out:
- Email: elvin@codingmindset.io
- X: @codingmindset
- GitHub Issues: For bug reports and feature requests
🙏 Acknowledgements
- Thanks to all our contributors and users for their valuable feedback and support.
- Special thanks to the open-source community for the amazing tools and libraries that made this project possible.
Built with ❤️ by Coding Mindset
Citing LLMDataLens
If you use LLMDataLens in your research, please cite it as follows:
@software{llmdatalens,
title = {LLMDataLens: A Framework for Evaluating LLM-based Applications},
author = {Elvin Gomez},
year = {2024},
url = {https://github.com/codingmindset/LLMDataLens.git},
}
Metadata
Release files for llm-data-lens 0.1.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_data_lens-0.1.5.tar.gz | 15.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_data_lens-0.1.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 33.6 kB
Release files / llm_data_lens-0.1.5.tar.gz
| Download URL | llm_data_lens-0.1.5.tar.gz |
|---|---|
| Size | 15.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
28f24322f7bb11b091c6a8d11d205aa6e195770f4849e2a5f527c5b759ffd5c9
|
|
BLAKE2b-256 checksum How to use checksums |
31300a8ce4712abe93a2996d75338d99d5096062f1bdbf32442085cd342cc53f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
poetry/1.8.3 CPython/3.10.5 Darwin/23.1.0
|
Release files / llm_data_lens-0.1.5-py3-none-any.whl
| Download URL | llm_data_lens-0.1.5-py3-none-any.whl |
|---|---|
| Size | 17.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5b937d9cd0b41c41af3c5d4614506b4a8becd6139aeeeaad8817e21b7b622e07
|
|
BLAKE2b-256 checksum How to use checksums |
889b4ccaf8c94a8a26ed0c8a7ffecaf00a217870ccefd1aee60c105475786322
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
poetry/1.8.3 CPython/3.10.5 Darwin/23.1.0
|