kenon
Semantic and co-occurrence graphs for midsized texts. Kenon builds weighted graphs from text using corpus-internal statistics — no neural models or external training data required. Supports co-occurrence windows, TF-IDF similarity, PMI embeddings, and disparity filter backbone extraction.
Installation
uv add kenon
python -m spacy download en_core_web_sm
Quickstart
from kenon import (
Tokenizer,
get_stopwords,
build_cooccurrence_graph,
extract_backbone,
)
# 1. Tokenize
tokenizer = Tokenizer("en_core_web_sm", lemmatize=True)
tokens = tokenizer.flat_tokens("The cat sat on the mat. The dog ran in the park.")
# 2. Build graph
stopwords = get_stopwords("english")
graph = build_cooccurrence_graph(tokens, window=2, stopwords=stopwords)
# 3. Extract backbone
backbone = extract_backbone(graph, min_alpha_ptile=0.3, min_degree=2)
print(f"Backbone: {backbone.number_of_nodes()} nodes, {backbone.number_of_edges()} edges")
Features
- Tokenization: spaCy-backed sentence splitting, tokenization, and lemmatization
- Stopwords: Merged NLTK + sklearn stopword lists with custom extensions
- Embeddings: Count vectors, TF-IDF, and PMI (via chronowords) — all corpus-internal
- Co-occurrence graphs: Skip-gram window co-occurrence with collocation detection
- Semantic graphs: Cosine similarity graphs from any embedder
- Backbone extraction: Disparity filter for statistically significant edges
Documentation
Full documentation — quickstart, the word-association tutorial, troubleshooting,
and the complete API reference — is at
kenon.readthedocs.io. The sources live in docs/.
Roadmap
Planned for the next iteration. The robustness items are analysed in detail in
PRE-MORTEM.md.
Robustness / API decisions
- Warn or document when
build_semantic_graphruns on too few documents — word-similarity is degenerate on small corpora (each word vector's dimension is the document count). - Avoid the dense
.toarray()materialisation in the sklearn embedders (out-of-memory risk on large corpora). - Make
transform()raise the documentedRuntimeErrorwhen called beforefit()(it currently surfaces sklearn'sNotFittedError). -
build_semantic_graph(k_neighbors=...): reuse the already-fitted embedder and warn instead of silently degrading to a threshold-only graph. - Guard
detect_collocationsagainst NLTK crashes on degenerate corpora —chi_sqraisesZeroDivisionErrorandlikelihooda math-domainValueErroron all-identical tokens; onlypmiis robust. - Wrap the unsupported-language error in
get_stopwordsand document its one-time NLTK download.
Proposed features
-
kenon.paths— a first-class concept-to-concept pathfinding helper (currently demonstrated via networkx in the tutorial). - Compare a text-derived network against human association norms — Nelson norms / Small World of Words (needs external datasets).
Made by
Kenon is made by Crow Intelligence.
License
MIT
Metadata
Release files for kenon 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| kenon-0.1.2.tar.gz | 14.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| kenon-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 31.9 kB
Release files / kenon-0.1.2.tar.gz
| Download URL | kenon-0.1.2.tar.gz |
|---|---|
| Size | 14.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e72313b1d3badc89de8ac046fe4b76b91921dfe6d93d7eefc7ad4e84af979169
|
|
BLAKE2b-256 checksum How to use checksums |
c6e7f6dcd036ced7ff57c7036f2f528ec282926eea06bc04aacaf63e538946c3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 29, 2026.
Transparency logRelease files / kenon-0.1.2-py3-none-any.whl
| Download URL | kenon-0.1.2-py3-none-any.whl |
|---|---|
| Size | 17.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2c88f2994a240fea523a3a40103c62d3eac046a81ca98ecf35725c20240fd428
|
|
BLAKE2b-256 checksum How to use checksums |
c7a47de6746d1bdb36dd4df4e6a511ae479a042b9e3c059f168490358317bd1a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 29, 2026.
Transparency log