Scalable Understanding of Datasets and Models with the Help of Large Language Models
I will make a video tutorial on this topic; stay tuned. This is the library and notebooks to help the audience understand my tutorial.
Installation
I recommend creating a conda environment with python >= 3.9 to use this package.
- Set up your openai key in your environment. i.e.
export OPENAI_API_KEY="[Your OPENAI API KEY]" - Installation
Option 1: clone and install locally (for developers)
git clone git@github.com:ruiqi-zhong/llm_explain.gitcd llm_explainpip install -e .
Option 2: install from github repo
pip3 install --no-cache-dir -v git+https://github.com/ruiqi-zhong/llm_explain.git
Option 3: install from pypi
pip3 install llm-explain
Usage
This repo supports the bare bone implementation for explaining dataset differences and clusters.
See the notebooks, llm_explain/tests/test_cluster.py, and llm_explain/tests/test_diff.py to understand how to use the functions implemented in this repo.
If you want to build on it, refer to other test files to understand the rest of the repo.
A quick example after installation
run python:
>>> from llm_explain.models.diff import explain_diff
>>> explain_diff(["cat", "dog", "fish", "carrot", "potato", "apple"], [False, False, False, True, True, True], proposer_num_rounds=2, proposer_num_explanations_per_round=2)
You will get outputs similar to the following in fewer than 30 seconds:
Printing top 3 explanations:
Explanation: refers to a plant-based item; specifically, the text mentions items that grow from plants, including vegetables and fruits. For example, 'carrot' is a type of root vegetable.
Accuracy: 1.0
Explanation: is a type of food; specifically, the text refers to items commonly recognized as food, typically vegetables or fruits. For example, 'apple' is known to be a fruit consumed as food.
Accuracy: 0.8333333333333333
Explanation: mentions a type of food; specifically, the text refers to something that is commonly eaten by humans. For example, 'This apple is very juicy.'
Accuracy: 0.8333333333333333
Notebooks
The notebooks illustrate the following sections in the video tutorial.
- 1.1 Core method: the proposer-validator framework
- 1.1 Extension 1: precise explanations.
- 1.1 Extension 2: goal-constrained explanations.
- 1.1 Extension 3: multiple explanations
- 1.2: Explainable clustering
Related works
related/references.pdf contains the related works mentioned in our presentation. related/main.tex contains the latex source file and references.bib contains the bibtex citations.
Metadata
Release files for llm-explain 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_explain-0.1.1.tar.gz | 16.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_explain-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 34.5 kB
Release files / llm_explain-0.1.1.tar.gz
| Download URL | llm_explain-0.1.1.tar.gz |
|---|---|
| Size | 16.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
531b00c410561884d05b20ccac662a2ef31017f1360c3e83f174aa1b14c426de
|
|
BLAKE2b-256 checksum How to use checksums |
2b544856baf79012df57b1f1e573181bc21d98e677839aca16cee8ea924da0ff
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.9.21
|
Release files / llm_explain-0.1.1-py3-none-any.whl
| Download URL | llm_explain-0.1.1-py3-none-any.whl |
|---|---|
| Size | 18.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
779b003b6103a39ee528563cab6f3215b11219dfd911a8e075de835608f7c332
|
|
BLAKE2b-256 checksum How to use checksums |
97bfee498511007aa8207b7555acf6f7fa251b4c585acf770f2815c1edfe8a05
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.9.21
|