20 Questions Game using LLMs
Project description
TwentyQ: The 20 Questions Game Engine (Python Package)
This pip package enables users to simulate and evaluate a 20 Questions-style guessing game using Large Language Models (LLMs) such as OpenAI GPT, Anthropic Claude, Google Gemini, and Hugging Face-hosted models. This package forms the backbone of the experiments in the paper: The World According to LLMs: How Geographic Origin Influences LLMs' Entity Deduction Capabilities
File Structure
.
├── LICENSE # MIT License for the package
├── pyproject.toml # Configuration file for pip installation
├── README.md # Documentation for usage
└── twentyq/ # Core package directory
├── __init__.py # Exports play_game for public use and loads environment variables
├── play.py # Main logic for running and saving the game
├── game.py # Game logic, turn simulation, message storage
├── utils.py # Helper functions for model loading and utilities
└── eval_result.py # Evaluation logic for computing metrics from saved game
Installation
pip install twentyq
Example Usage
from twentyq import play_game
result = play_game(
name="Lionel Messi",
guesser_model_name="gpt-4o", # Supports OpenAI, Claude, Gemini, Hugging Face
num_turns=20,
temperature=0.8,
language="english",
base_output_dir="./outputs",
game_type="celebrities" # "celebrities" or "things"
)
Input Parameters
name: (str) – The target entity (celebrity or thing) to be guessed (e.g., "Lionel Messi").guesser_model_name: (str) – Model to be used. Supported:- OpenAI models (e.g.,
gpt-4o) - Claude (e.g.,
claude-3-5-sonnet-latest) - Gemini (e.g.,
gemini-2.0-flash) - Hugging Face models (e.g.,
meta-llama/Llama-3.3-70B-Instruct)
- OpenAI models (e.g.,
num_turns: (int) – Number of allowed turns (e.g., 20).temperature: (float) – Sampling temperature for the guesser model (e.g., 0.8).language: (str) – The language to conduct the game in. Supported:
english,hindi,mandarin,spanish,japanese,turkish,frenchbase_output_dir: (str) – Base path for saving results.game_type: (str) –"celebrities"or"things"depending on the entity type.
API Keys
Regardless of which model is used, the environment must have the following keys set:
OPENAI_API_KEY– (mandatory)- If using Hugging Face models:
HUGGINGFACE_TOKEN - If using Gemini:
GENAI_API_KEY - If using Claude:
ANTHROPIC_API_KEY
You can either export these directly in your shell:
export OPENAI_API_KEY=your_key_here
export HUGGINGFACE_TOKEN=your_token_here
Or save them in a .env file in your project directory:
OPENAI_API_KEY=your_key_here
HUGGINGFACE_TOKEN=your_token_here
Output
- A folder named like
{guesser_model_name}_{num_turns}_turns_{language}will be created insidebase_output_dir(e.g.,gpt-4o_20_turns_english) - Inside this folder, a
.txtfile named after the entity (e.g.,Lionel Messi.txt) will contain the full game conversation.
./outputs/gpt-4o_20_turns_english/
├── Lionel Messi.txt # Full dialogue from the 20Q game
Return Value
The play_game function returns a dictionary like:
{
"Lionel Messi": {
"guesser": [
"Is the person a footballer?",
"Has he won a Ballon d'Or?",
"Is it Lionel Messi?"
],
"answerer": [
"Yes.",
"Yes.",
"Bingo!"
],
"Success": True,
"Has_Given_Up": False,
"Turn_At_Which_Given_Up": "N/A",
"Num_Turns_To_Answer": 3,
"Num_Yes": 3
}
}
If Has_Given_Up is True, then Num_Turns_To_Answer will be "N/A".
If Success is True, then Turn_At_Which_Given_Up will be "N/A".
Design Notes
To ensure consistency and mitigate knowledge disparities, the same model is used for both the Guesser and Judge roles.
We use a temperate of 0.2 for the Judge to ensure it provides consistent and deterministic responses.
We use max_tokens of 100 for the Guesser to allow for concise responses while still being able to ask follow-up questions
We use max_tokens of 30 for the Judge to ensure it just provides a simple "Yes", "No", or "Maybe/Dunno" response.
Citation
If you use this python package, please cite:
@inproceedings{lalai20q,
title={The World According to LLMs: How Geographic Origin Influences LLMs' Entity Deduction Capabilities},
author={Lalai, Harsh Nishant and Shah, Raj Sanjay and Pei, Jiaxin and Varma, Sashank and Wang, Yi-Chia and Emami, Ali},
booktitle={Second Conference on Language Modeling}
}
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file twentyq-0.0.7.tar.gz.
File metadata
- Download URL: twentyq-0.0.7.tar.gz
- Upload date:
- Size: 15.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ed3314dad12957bc699e7b5d22648b3a13311055d69574f8b787e9e8ad211617
|
|
| MD5 |
0b76b75a8a1ed30afe67daacf0a4bb66
|
|
| BLAKE2b-256 |
9f4da8dccafdf25d593197d2d096d8c5c3683ee70fcf2a08bbcb0a99cda2ab4a
|
File details
Details for the file twentyq-0.0.7-py3-none-any.whl.
File metadata
- Download URL: twentyq-0.0.7-py3-none-any.whl
- Upload date:
- Size: 14.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bf982ee37642f4c056e32308f343fd2455dd4bc1a98fb9d15f185cf6a3ff1511
|
|
| MD5 |
6e69a790738b0f78263458afbcf55a7c
|
|
| BLAKE2b-256 |
9a5a3064a252ba51a41845bdab2f41f6df5346098a774c9222ffd94cc015e28f
|