Skip to main content

WordTangible

PyPI - Version DOI

WordTangible is a simple Python library for analyzing the concreteness and imageability of words and text. This can be useful for various natural language processing tasks, readability analysis, and linguistic research.

Data is pulled from:

  • The Brysbaert dataset (Brysbaert et al., 2014)[1]
  • The Glasgow dataset (Scott et al., 2019)[2]
  • The MRC Psycholinguistic Database (Coltheart, 1981; machine-usable dictionary: Wilson, 1988)[3][4]

The default rating is a quality-ordered fallback on a 1-5 scale (5 = most concrete): Brysbaert's raw value when a word is in Brysbaert (the largest and most recent source, natively 1-5), otherwise Glasgow, otherwise MRC, each linearly rescaled to 1-5. The sources are deliberately not averaged — their normalized distributions have systematically different means, so a linear-rescale average would skew multi-source words rather than reduce noise. This way every value is a real published rating from a single identifiable study (and ~99% of words return Brysbaert's exact published value). The raw per-source ratings, an MRC-free variant, and a three-way mean are all available via the source parameter (see below).

Features

  • Get concreteness ratings for individual words
  • Choose the ratings source: the default fallback, any single dataset un-normalized (brysbaert, glasgow, mrc), an MRC-free variant (open) for commercial use, or a normalized three-way mean
  • Rated two-word compounds match as units: "baseball bat" gets the compound's own rating (Brysbaert rated 2,896 such expressions) instead of blending "baseball" and "bat"
  • Unrated inflected forms score as their WordNet lemma ("whales" as "whale", "replied" as "reply") — worth ~10 points of token coverage on typical fiction; disable with lemma_fallback=False for values strictly comparable to the published norms
  • Calculate average concreteness for a given text
  • Compute the ratio of concrete to abstract words in a text (with optional add-k smoothing to keep it finite and stable on short texts)
  • Report rating coverage — how much of a text the concreteness mean actually rests on

Installation

You can install WordTangible using pip:

pip install wordtangible

Usage

Here are some basic examples of how to use WordTangible:

from wordtangible import word_concreteness, avg_text_concreteness, concrete_abstract_ratio

# Get concreteness rating for a single word
print(word_concreteness("apple"))  # Output: 5.0 (highly concrete)

# Calculate average concreteness of a text
text = "The abstract concept of love is as tangible as the apple in your hand."
print(avg_text_concreteness(text))  # Output: ~2.9 (mix of concrete and abstract)

# Get the ratio of concrete to abstract words
print(concrete_abstract_ratio(text))  # Output: ~1.0 (balanced concrete and abstract words)

# Smoothed ratio: finite even when a text has no very-abstract words,
# and steadier on short texts (add-k smoothing; default k=0 keeps the
# classic behavior, including float('inf') for the no-abstract case)
print(concrete_abstract_ratio(text, smoothing=1))

# How much of the text the concreteness mean rests on (0.0-1.0):
# a mean at 0.85 coverage is trustworthy; the same mean at 0.12
# (jargon, dialect, names) is noise
from wordtangible import concreteness_coverage
print(concreteness_coverage(text))

Choosing a ratings source

Every function accepts a source parameter:

word_concreteness("apple")               # 5.0   — default fallback, 1-5 scale
word_concreteness("apple", "brysbaert")  # 5.0   — raw Brysbaert, 1-5 scale
word_concreteness("apple", "glasgow")    # 6.824 — raw Glasgow CNC, 1-7 scale
word_concreteness("apple", "mrc")        # 620   — raw MRC CNC, 100-700 scale
word_concreteness("apple", "open")       # 5.0   — like default, but never MRC
word_concreteness("apple", "mean")       # 4.78  — mean of all three, rescaled to 1-5

avg_text_concreteness(text, source="open")
  • default — Brysbaert, else Glasgow, else MRC (rescaled to 1-5).
  • brysbaert / glasgow / mrc — one dataset's raw, un-normalized values on its native scale; None for words it doesn't rate. Useful for comparing directly against the published norms. If you pass these to concrete_abstract_ratio, adjust its thresholds to the source's scale.
  • open — Brysbaert, else Glasgow: excludes the MRC database, whose terms are "for research purposes" (see licensing below), making this the right choice for commercial products.
  • mean — the mean of whichever of the three rate the word, each linearly rescaled to 1-5. Beware: linear rescaling doesn't fully align the scales (the sources' normalized means differ systematically), so these values aren't comparable to any single set of published norms.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

The WordTangible code is licensed under the MIT License - see the LICENSE file for details.

Data sources & licensing

The bundled ratings file (wordtangible/resources/concreteness_ratings.csv) is derived from third-party datasets, and the MIT license above does not apply to that data. Each source has its own terms:

  • Glasgow Norms (Scott et al., 2019): published open access under a Creative Commons Attribution 4.0 license — redistribution of derived data is permitted with attribution, which the citation below provides.
  • Brysbaert concreteness norms (Brysbaert et al., 2014): made freely available by the authors (via the Behavior Research Methods supplement and the Ghent CRR lab) and widely redistributed in research software; no formal license accompanies the data, so provenance and citation are provided here as is standard practice.
  • MRC Psycholinguistic Database (Coltheart, 1981; Wilson, 1988): the database's distribution terms state that it is available for research purposes. The default ratings fall back to MRC-derived values for the few hundred words neither Brysbaert nor Glasgow rates. If you intend to use WordTangible in a commercial (non-research) product, pass source="open" — it draws only on Brysbaert and Glasgow — or verify the MRC terms for your use case.

The bundled CSV keeps each source's raw rating in its own column, and scripts/build_ratings.py regenerates it from the original datasets (downloaded on demand; the raw files are not stored in this repository).

This section documents provenance in good faith and is not legal advice.

Citing WordTangible

If you use WordTangible in research, please cite both the tool and the rating datasets your results rest on (all three under the default source; Brysbaert and Glasgow only if you use source="open"; the single dataset if you use a raw source). GitHub's "Cite this repository" button generates a citation from CITATION.cff, or use:

Robison, J. (2026). WordTangible (Version 0.5.0) [Computer software]. https://github.com/jrrobison1/wordtangible

@software{robison_wordtangible,
  author  = {Robison, Jason},
  title   = {WordTangible},
  version = {0.5.0},
  year    = {2026},
  url     = {https://github.com/jrrobison1/wordtangible}
}

Dataset citations are given in full in the References below.

References

[1] Brysbaert, M., Warriner, A. B., & Kuperman, V. (2014). Concreteness ratings for 40 thousand generally known English word lemmas. Behavior Research Methods, 46(3), 904-911. https://doi.org/10.3758/s13428-013-0403-5

[2] Scott, G. G., Keitel, A., Becirspahic, M., Yao, B., & Sereno, S. C. (2019). The Glasgow Norms: Ratings of 5,500 words on nine scales. Behavior Research Methods, 51(3), 1258-1270. https://doi.org/10.3758/s13428-018-1099-3

[3] Coltheart, M. (1981). The MRC psycholinguistic database. The Quarterly Journal of Experimental Psychology Section A, 33(4), 497-505. https://doi.org/10.1080/14640748108400805

[4] Wilson, M. (1988). MRC Psycholinguistic Database: Machine-usable dictionary, version 2.00. Behavior Research Methods, Instruments, & Computers, 20(1), 6-10. https://doi.org/10.3758/BF03202594

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

wordtangible-0.5.0.tar.gz (277.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

wordtangible-0.5.0-py3-none-any.whl (278.9 kB view details)

Uploaded Python 3

File details

Details for the file wordtangible-0.5.0.tar.gz.

File metadata

  • Download URL: wordtangible-0.5.0.tar.gz
  • Upload date:
  • Size: 277.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for wordtangible-0.5.0.tar.gz
Algorithm Hash digest
SHA256 55caf6f0cfc59e493d53fb7460b0836cad14c773bfcf5301ed299438f3e6c189
MD5 79374b7bff249b0935c20c31e7e412c9
BLAKE2b-256 f882d40f851c024efaad35d7325b1231ffd711676ee3529f7c1fca82a7fea34b

See more details on using hashes here.

Provenance

The following attestation bundles were made for wordtangible-0.5.0.tar.gz:

Publisher: publish-to-pypi.yml on jrrobison1/wordtangible

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file wordtangible-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: wordtangible-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 278.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for wordtangible-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 3af604952756a3408a813f3362cd93ae952c183d7eb65ab7c8e60d9f26ecc0c7
MD5 eaf7648db4ab4d247cab565ed6766e17
BLAKE2b-256 49fb81ee5869f91f67b72bb3ea9a9892f5db229b56090595ac1030c9663fed96

See more details on using hashes here.

Provenance

The following attestation bundles were made for wordtangible-0.5.0-py3-none-any.whl:

Publisher: publish-to-pypi.yml on jrrobison1/wordtangible

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.6.0

2 files

This release

0.5.0 This release

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page