Skip to main content

R.U.Psycho

CI Docs Python License: MIT arXiv

R.U.Psycho (Robust Unified Psychometric Testing of Language Models) is a framework for designing and running robust and reproducible psychometric experiments on generative language models, with limited coding expertise required.

Documentation: https://julianschelb.github.io/rupsycho/

Paper: R.U.Psycho? Robust Unified Psychometric Testing of Language Models (Schelb, Borin, Garcia & Spitz, 2025)

Why?

Instabilities in model outputs, sensitivity to prompt design and generation parameters, and the sheer number of model versions make psychometric studies of language models hard to reproduce. R.U.Psycho turns a whole study into one declarative configuration: the questionnaire, the personas the model answers as, the models and their parameters, the prompt template and the random seeds. The package runs every model × seed × persona × item combination, turns the free-text answers into scorable responses and computes item and scale scores.

Features

  • Declarative experiments in a single, shareable JSON file (secrets are never exported)
  • Seeds that work: every back-end receives the seed in the way it supports (see Reproducibility)
  • Many back-ends: local & remote Hugging Face, Ollama, OpenAI, Google, DeepSeek, any LangChain runnable
  • Robust runs: a RunSummary instead of silent failures, an error policy, opt-in concurrency for API models with a deterministic result order, re-runnable experiments
  • Post-processing and scoring: cleaners, refusal / "as an AI" validators, rule- and model-based judges, then weights, reverse-keyed items and per-trait scale scores
  • Command line: rupsycho run | validate | prompt | postprocess | examples | configurator
  • Configurator app with LLM-assisted questionnaire import from PDF
  • Light core: import rupsycho takes milliseconds; model back-ends are optional extras

Installation

pip install git+https://github.com/julianschelb/rupsycho.git          # core: configs, run loop, parsers, scoring
pip install "rupsycho[huggingface] @ git+https://github.com/julianschelb/rupsycho.git"   # + local Hugging Face models

Extras: huggingface (PyTorch, Transformers), openai, ollama, google, deepseek, models (all back-ends), configurator (Streamlit app), notebook, quantization, all. Using a back-end whose extra is missing raises an error that names the extra to install.

Requires Python 3.10 or newer (tested on 3.10 – 3.14).

Quick start

import rupsycho as rup

# A bundled example: five Big Five items answered by two personas
experiment = rup.load_example_experiment("bfi", seeds=[1, 2, 3])
experiment.print_assembled_prompt(item_idx=1, persona_idx=0)   # exactly what the model sees

summary = experiment.run()                      # needs the huggingface extra for the example model
print(summary)                                  # e.g. "30 model calls in 41.2s"
answers = experiment.get_answers_as_dataframe() # one row per model x seed x persona x item

Use your own model instead of the example's:

from langchain_openai import ChatOpenAI          # pip install "rupsycho[openai]"

experiment = rup.load_example_experiment("bfi", models={}, seeds=[1, 2, 3])
experiment.add_model(ChatOpenAI(model="gpt-4o-mini"), identifier="gpt-4o-mini")
experiment.run(max_concurrency=8)                # API calls run in parallel, order is preserved

From the shell:

rupsycho examples copy bfi bfi.json
rupsycho validate bfi.json
rupsycho run bfi.json -o results.csv --seeds 1 2 3
rupsycho postprocess bfi.json results.csv -o processed.csv

Then score the judged answers:

import pandas as pd
from rupsycho import scoring

processed = pd.read_csv("processed.csv")
scored = scoring.score_answers(processed, experiment)
print(scoring.scale_scores(scored))              # mean score per model, persona, seed and trait

See the Getting Started guide, the tutorials and the notebooks in examples/.

Development

pip install -e ".[dev,models,configurator]"
pre-commit install --hook-type pre-commit --hook-type commit-msg

poe check              # lint + format check + mypy + offline tests
poe docs               # serve the docs locally
poe test-integration   # tests that download real models

See CONTRIBUTING.md and the development guide.

Citation

If you use R.U.Psycho, please cite the paper (see also CITATION.cff):

@misc{schelb2025rupsycho,
  title         = {R.U.Psycho? Robust Unified Psychometric Testing of Language Models},
  author        = {Schelb, Julian and Borin, Orr and Garcia, David and Spitz, Andreas},
  year          = {2025},
  eprint        = {2503.10229},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2503.10229}
}

License

MIT

Metadata

Release files for rupsycho 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rupsycho 1.0.0
File Size Uploaded
rupsycho-1.0.0.tar.gz 89.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rupsycho 1.0.0
File Interpreter ABI Platform
rupsycho-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 194.4 kB

Release files / rupsycho-1.0.0.tar.gz

Download URL rupsycho-1.0.0.tar.gz
Size 89.2 kB
Tags Source
SHA-256 checksum
How to use checksums
a010898cf69039ae94ec70a81b1a742016d4629f7731c9d545f784997c171dfd
BLAKE2b-256 checksum
How to use checksums
2115bd0d7f9c2b19f311c4a6bd1ade5a05df30e47766f8384a763d63fbc6bbc7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / rupsycho-1.0.0-py3-none-any.whl

Download URL rupsycho-1.0.0-py3-none-any.whl
Size 105.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9164a7b41d572b7cc2f15e0f5ab5949d455ea0f090ba6eaacfbf407bb2f5c3f3
BLAKE2b-256 checksum
How to use checksums
97c2f880fa7130169ccc69c9fe8301993f3b6b699760a95acc4f9794592d3dda
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page