Skip to main content

1Q1W (One Question One World)

Generate a set of unique evaluation criteria from one question or a list of questions using LLMs. The pipeline expands scenarios, perspectives, and criteria through multiple iterations, then deduplicates and assigns scores.

Installation

From PyPI:

pip install oneq1w

From GitHub:

pip install "oneq1w @ git+https://github.com/mims-harvard/1Q1W.git"

From local clone (for development):

git clone https://github.com/mims-harvard/1Q1W.git
cd 1Q1W
pip install -e .

API Keys (.env)

Create a .env file in the project root with the keys for your chosen provider:

# OpenAI (default)
OPENAI_API_KEY=sk-...

# Azure OpenAI
AZURE_OPENAI_API_KEY=...
AZURE_OPENAI_ENDPOINT=https://your-resource.openai.azure.com

# Claude
ANTHROPIC_API_KEY=sk-ant-...

# Gemini
GOOGLE_API_KEY=...   # or GEMINI_API_KEY

# Grok
XAI_API_KEY=...

# DeepSeek
DEEPSEEK_API_KEY=...

# vLLM
VLLM_SERVER_URL=http://localhost:8000/v1

Load API keys before running or use python-dotenv to load .env in your script.

Quick Start

from oneq1w import CriteriaGenerator

gen = CriteriaGenerator(model="gpt-4o")

# Single question (string)
result = gen.generate("What is machine learning?")
print(result["final_criteria"])

# Single question (dict)
result = gen.generate({"id": "q1", "question": "What is AI?"})

# Batch
results = gen.generate([
    {"id": "q1", "question": "What is AI?"},
    {"id": "q2", "question": "How does deep learning work?"},
])

Input Arguments

1. CriteriaGenerator Init Parameters

Parameter Type Default Description
model str "gpt-4.1" Model name (auto-detects provider)
base_url str None API base URL (required for vLLM)
api_key str None API key (uses env vars if not provided)
temperature float 0.4 Generation temperature
embedding_model str "text-embedding-3-small" Model for embeddings (deduplication)
n_scenario_expands int 3 Scenario expansion iterations
n_perspective_expands int 4 Perspective expansion iterations
n_criteria_expands int 3 Criteria expansion iterations
dedup_threshold float 0.6 Cosine similarity threshold for deduplication (0-1)
max_workers int 8 Parallel workers for batch processing
max_retries int 5 Max retries on rate limit errors
debug bool False Print raw LLM outputs for debugging

2. generate() Input Arguments

Input Type Format Required Optional
str "question text" — —
dict {"id": "...", "question": "..."} id, question image, web_content
list [{...}, {...}] each item has id, question image, web_content per item
  • id: Identifier for the question (auto-assigned if omitted).
  • question: The question text.
  • image: Base64 string (with or without data:image/...;base64, prefix) for vision-capable models.
  • web_content: Retrieved web context; appended to the question as [Retrieved Web Context] ... [End of Web Context].

Output Format

{
    "id": "q1",
    "question": "What is AI?",
    "scenarios": [...],
    "raw_perspectives": [...],
    "reviewed_perspectives": [...],
    "raw_criteria": [...],
    "reviewed_criteria": [...],
    "final_criteria": [
        {"criterion": "Explains core concepts", "points": 3},
        {"criterion": "Provides examples", "points": 2},
    ],
}

Supported Models

Provider Models Env Variable
OpenAI gpt-5, gpt-4.1, ... OPENAI_API_KEY
Azure OpenAI gpt-5, gpt-4.1, ... AZURE_OPENAI_API_KEY, AZURE_OPENAI_ENDPOINT
Claude claude-4-5-opus, claude-4-5-sonnet, ... ANTHROPIC_API_KEY
Gemini gemini-3-pro, gemini-3-flash, ... GOOGLE_API_KEY
Grok grok-4.1-fast, ... XAI_API_KEY
DeepSeek deepseek-chat, deepseek-reasoner DEEPSEEK_API_KEY
vLLM Qwen3-30B, ... VLLM_SERVER_URL or base_url param

vLLM: Uses a separate server for hosting. You need to install the vLLM package suitable for your server environment or use the official vLLM Docker image. Once the server is running, provide the model name and the server URL (e.g. VLLM_SERVER_URL=http://localhost:8000/v1 or base_url when initializing CriteriaGenerator).

Data

Raw data and generated criteria (gpt-4.1 with scenario expand x3, perspective expand x4, criteria expand x3) are available at https://huggingface.co/datasets/suyc21/1Q1W.

Test Scripts

This folder contains several test scripts:

Script Description
test_providers.py Tests different providers (OpenAI, Claude, Gemini, Grok, DeepSeek, vLLM) on HealthBench sample data or HLE sample data. Usage: python test_providers.py --provider openai
test_pipeline.py Tests generating HealthBench data using gpt-4.1.
evaluate_criteria.py Evaluates pre-generated criteria against expert rubrics. Computes coverage and uniqueness metrics. Use --analyze-only to compute scores only without re-running LLM evaluation. Usage: python evaluate_criteria.py -i results.json --analyze-only

Release files for oneq1w 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for oneq1w 0.1.2
File Size Uploaded
oneq1w-0.1.2.tar.gz 21.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for oneq1w 0.1.2
File Interpreter ABI Platform
oneq1w-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 43.6 kB

Release files / oneq1w-0.1.2.tar.gz

Download URL oneq1w-0.1.2.tar.gz
Size 21.9 kB
Tags Source
SHA-256 checksum
How to use checksums
b402718ca63d15d4e48623f3db6b81b0632baae142fe733d93aeb1a526d4358c
BLAKE2b-256 checksum
How to use checksums
5d90047a700f6866b27a4736f4e191b9648cd40b82d27ecb3483f28cd7f499a3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.14

Release files / oneq1w-0.1.2-py3-none-any.whl

Download URL oneq1w-0.1.2-py3-none-any.whl
Size 21.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
527781a542403d910f65758d68cda07da8a061b10dabb5cceea12aa2d0fa7789
BLAKE2b-256 checksum
How to use checksums
2bd264c59690c3c7bed4f5ca12a663723d4492f81b82cd92cb566c75fa7a224b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.14

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page