Skip to main content

Generate evaluation criteria for any question using LLMs

Project description

Qworld

Generate a set of unique evaluation criteria from one question or a list of questions using LLMs. The pipeline expands scenarios, perspectives, and criteria through multiple iterations, then deduplicates and assigns scores.

Installation

From PyPI:

pip install qworld

From PyPI with local embedding fallback (optional extra):

pip install "qworld[local-embeddings]"

From GitHub:

pip install "qworld @ git+https://github.com/mims-harvard/Qworld.git"

From local clone (for development):

git clone https://github.com/mims-harvard/Qworld.git
cd Qworld
pip install -e .

Embedding Dependency Behavior

  • If OPENAI_API_KEY or AZURE_OPENAI_API_KEY is available, qworld uses OpenAI/Azure embeddings.
  • If GOOGLE_API_KEY is available, qworld uses Gemini embeddings.
  • If non of these keys are available, qworld falls back to local sentence-transformers embeddings.
  • To enable that fallback, install the optional extra: pip install "qworld[local-embeddings]".

API Keys (.env)

Create a .env file in the project root with the keys for your chosen provider:

# OpenAI (default)
OPENAI_API_KEY=sk-...

# Azure OpenAI
AZURE_OPENAI_API_KEY=...
AZURE_OPENAI_ENDPOINT=https://your-resource.openai.azure.com

# Claude
ANTHROPIC_API_KEY=sk-ant-...

# Gemini
GOOGLE_API_KEY=...   # or GEMINI_API_KEY

# Grok
XAI_API_KEY=...

# DeepSeek
DEEPSEEK_API_KEY=...

# vLLM
VLLM_SERVER_URL=http://localhost:8000/v1

Load API keys before running or use python-dotenv to load .env in your script.

Quick Start

from qworld import CriteriaGenerator

gen = CriteriaGenerator(model="gpt-4.1")

# Single question (string)
result = gen.generate("What is machine learning?")
print(result["final_criteria"])

# Single question (dict)
result = gen.generate({"id": "q1", "question": "What is AI?"})

# Batch
results = gen.generate([
    {"id": "q1", "question": "What is AI?"},
    {"id": "q2", "question": "How does deep learning work?"},
])

Input Arguments

1. CriteriaGenerator Init Parameters

Parameter Type Default Description
model str "gpt-4.1" Model name (auto-detects provider)
base_url str None API base URL (required for vLLM)
api_key str None API key (uses env vars if not provided)
temperature float 0.4 Generation temperature
embedding_model str "text-embedding-3-small" Model for embeddings (deduplication)
n_scenario_expands int 3 Scenario expansion iterations
n_perspective_expands int 4 Perspective expansion iterations
n_criteria_expands int 3 Criteria expansion iterations
dedup_threshold float 0.7 Cosine similarity threshold for deduplication (0-1)
max_workers int 8 Parallel workers for batch processing
max_retries int 5 Max retries on rate limit errors
debug bool False Print raw LLM outputs for debugging

2. generate() Input Arguments

Input Type Format Required Optional
str "question text"
dict {"id": "...", "question": "..."} id, question image, web_content
list [{...}, {...}] each item has id, question image, web_content per item
  • id: Identifier for the question (auto-assigned if omitted).
  • question: The question text.
  • image: Base64 string (with or without data:image/...;base64, prefix) for vision-capable models.
  • web_content: Retrieved web context; appended to the question as [Retrieved Web Context] ... [End of Web Context].

Output Format

{
    "id": "q1",
    "question": "What is AI?",
    "scenarios": [...],
    "raw_perspectives": [...],
    "reviewed_perspectives": [...],
    "raw_criteria": [...],
    "reviewed_criteria": [...],
    "final_criteria": [
        {"criterion": "Explains core concepts", "points": 3},
        {"criterion": "Provides examples", "points": 2},
    ],
}

Supported Models

Provider Models Env Variable
OpenAI gpt-5, gpt-4.1, ... OPENAI_API_KEY
Azure OpenAI gpt-5, gpt-4.1, ... AZURE_OPENAI_API_KEY, AZURE_OPENAI_ENDPOINT
Claude claude-4-5-opus, claude-4-5-sonnet, ... ANTHROPIC_API_KEY
Gemini gemini-3-pro, gemini-3-flash, ... GOOGLE_API_KEY
Grok grok-4.1-fast, ... XAI_API_KEY
DeepSeek deepseek-chat, deepseek-reasoner DEEPSEEK_API_KEY
vLLM Qwen3-30B, ... VLLM_SERVER_URL or base_url param

vLLM: Uses a separate server for hosting. You need to install the vLLM package suitable for your server environment or use the official vLLM Docker image. Once the server is running, provide the model name and the server URL (e.g. VLLM_SERVER_URL=http://localhost:8000/v1 or base_url when initializing CriteriaGenerator).

Data

Raw data and generated criteria (gpt-4.1 with scenario expand x3, perspective expand x4, criteria expand x3) are available at https://huggingface.co/datasets/suyc21/qworld.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

qworld-0.1.1.tar.gz (22.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

qworld-0.1.1-py3-none-any.whl (22.0 kB view details)

Uploaded Python 3

File details

Details for the file qworld-0.1.1.tar.gz.

File metadata

  • Download URL: qworld-0.1.1.tar.gz
  • Upload date:
  • Size: 22.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.13

File hashes

Hashes for qworld-0.1.1.tar.gz
Algorithm Hash digest
SHA256 31725313f9a2f83efcef734ab53fc9bc2def5209ee2d093e99f6ffc41c6520db
MD5 3813e680dc5df6fb9b7d2592868f3186
BLAKE2b-256 4d6f4627e676ff17c39a3debae277088112cccbc925bae0ce15c0538909760ef

See more details on using hashes here.

File details

Details for the file qworld-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: qworld-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 22.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.13

File hashes

Hashes for qworld-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 705906f184d1e204d6335e1b7de90878377999da0c302cd20ebe75d05b9ed5db
MD5 12e5be9d11b0338bd4c787436e27f743
BLAKE2b-256 c26d44fb49bc9f61b8c37642da041c5778d3ca7af7b5dee1d9787d369f96f499

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page