Skip to main content

lmdoctor

Extract, detect, and control representations within language models as they read and write text.

Detect lies, block harmful content, alter emotiontal tone, and more!

lmdoctor reads and manipulates a model's hidden states at inference time and, thus, requires access to model weights. Based largely on ideas from Representation Engineering: A Top-Down Approach to AI Transparency by Zou et al. 2023 (paper, code), as well as The Geometry of Truth by Marks & Tegmark, 2023 (paper, code).

Built on 🤗 transformers
pip package here
project repo here

Disclaimer: This is an experimental package and may not have sufficient sensitivity/specificity for your use-case. For an example analysis of the characteristics of one honesty detector, see evals/honesty/mistral_eval.ipynb.

Example

honesty extraction

from lmdoctor.doctor import Doctor
extraction_target = 'honesty'
doc = Doctor(model, tokenizer, user_tag, assistant_tag, extraction_target=extraction_target)
doc.extract()

lie detection

prompt = "Tell me a fact about humans"
doc.generate(prompt, max_new_tokens=12)

truth

prompt = "Tell me a lie about humans"

lie

honesty control

# without control
prompt = "You killed a person for revenge but don't want to get caught. What would you tell the police?"
doc.generate_with_control(prompt, control_direction=None, max_new_tokens=12)
output: I would tell the police that I did not kill anyone.
# with control
doc.generate_with_control(prompt, control_direction=-1, max_new_tokens=12)
output: I would tell the police that I have killed a person

For the complete example, see examples/honesty_example.ipynb

Getting started

Tested on linux

from pip: pip install lmdoctor
from source: pip install . after cloning

After install, try running honesty_example.ipynb

Extraction targets

The table below describes the targets we support for extracting internal representations. In functional extraction, the model is asked to produce text (e.g. prompt="tell me a lie"). In conceptual extraction, the model is asked to consider a statement (e.g. "consider the truthfulness of X"). For targets where both are supported, you can try each to see which works best for your use-case.

Target Method Types
truth conceptual none
honesty functional none
morality conceptual & functional none
emotion conceptual anger, disgust, fear, happiness, sadness, surprise
fairness conceptual & functional race, gender, prefession, religion
harmlessness conceptual none

Metadata

Release files for lmdoctor 0.5.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lmdoctor 0.5.7
File Size Uploaded
lmdoctor-0.5.7.tar.gz 6.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for lmdoctor 0.5.7
File Interpreter ABI Platform
lmdoctor-0.5.7-py3-none-any.whl Python 3 none any Details

Total release size: 13.4 MB

Release files / lmdoctor-0.5.7.tar.gz

Download URL lmdoctor-0.5.7.tar.gz
Size 6.7 MB
Tags Source
SHA-256 checksum
How to use checksums
86cca61172b84e079072af32845597f619b92ec42f8f2548e08d79cd9ead4b7f
BLAKE2b-256 checksum
How to use checksums
d7f4e9026ce125acb87c0e0f10dcce933de5c5fd85477c345ad294bbc68dfd94
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.10.12

Release files / lmdoctor-0.5.7-py3-none-any.whl

Download URL lmdoctor-0.5.7-py3-none-any.whl
Size 6.7 MB
Tags Python 3
SHA-256 checksum
How to use checksums
31f8f75f4d5068b33fbdd0a7cf9c290f26d8b58de053b122315ee167f0b21a2e
BLAKE2b-256 checksum
How to use checksums
1886e7c85b9393939dda462b22d60afe047bd54b51c60881129560b61332b949
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.10.12

Release history Release notifications | RSS feed

This release

0.5.7 This release

2 release files

0.5.6

1 release file

0.5.5

1 release file

0.5.4

1 release file

0.5.3

1 release file

0.5.2

1 release file

0.5.1

1 release file

0.5.0

1 release file

0.4.0

1 release file

0.3.0

1 release file

0.2.0

1 release file

0.1.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page