SABER
SABER (Sistema de Análise e Busca de Estruturas Relevantes) identifies European Portuguese grammatical structures in text and reports the CEFR level (A1–C2) at which each one is taught, following the Referencial Camões.
It gives you two things:
- Spans — every occurrence of every structure, with its text, its token and character offsets, and its CEFR level. One row per occurrence.
- Features — per document, the raw frequency, the frequency per 100 tokens, and the presence (0/1) of each structure. One row per document, ready to feed into a model or a statistical analysis.
252 structures are covered, described in Portuguese and organised into 11 broad groups (nouns, adjectives, verbs, adverbs, pronouns, determiners, quantifiers, relations between constituents, sentence types, sentence polarity, relations between sentences) and 31 finer categories. You can extract all of them, or filter by group, category, CEFR level, or individual structure.
Table of contents
- Installation
- Quick start
- Input: strings, files, folders
- Choosing which structures to extract
- Output columns
- How it works
- Performance
- Working on the matchers
- Testing the matchers
- Licence and attribution
Installation
Python 3.9 or newer, but not above 3.13.
From a clone:
git clone https://github.com/sorooshakef/SABER-Toolkit.git
cd SABER-Toolkit
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e .
Then download the Portuguese models once per machine (~1 GB, mostly Stanza):
python -c "import saber; saber.download_models()"
Note on PyTorch.
torchis pinned to 2.5.1 because Stanza's model loading breaks with newer releases. If another package upgrades it, reinstall withpip install torch==2.5.1.
Quick start
import saber
spans = saber.extract("Ele foi ao cinema no domingo com a minha irmã.", mode="spans")
print(spans[["structure", "level", "text", "description"]])
structure level text description
0 a3d1_12_A2 A2 foi Pretérito perfeito simples do indicativo - ver...
1 a6d1_6_A1 A1 ao cinema Artigo definido - concordância - em género e n...
2 a1d1_7_A2 A2 cinema Nomes masculinos terminados em "a"
3 a6d1_10_A2 A2 no Artigo definido - contração com prep...
4 a6d1_6_A1 A1 no domingo Artigo definido - concordância - em género e n...
5 a6d1_5_A1 A1 a minha Artigo definido - uso/valor - antes de determi...
6 a6d1_6_A1 A1 a minha irmã Artigo definido - concordância - em género e n...
7 a5d3_1_A1 A1 minha irmã Pronomes Possessivos - variação em pessoa, gén...
And the feature matrix for a folder of texts:
from pathlib import Path
features = saber.extract(Path("corpus/"), mode="features")
features.to_csv("features.csv", index=False)
extract always returns a pandas.DataFrame. saber.extract_spans(...) and
saber.extract_features(...) are aliases for the two modes.
Input: strings, files, folders
The first argument accepts text, a file, a folder, or a mix of them. The distinction between "a string of text" and "a path" is made by type, so it is never guessed:
saber.extract("Um texto qualquer.") # str -> raw text
saber.extract(Path("texto.txt")) # Path -> one file
saber.extract(Path("corpus/")) # Path -> every text file in the folder
saber.extract(["Primeiro texto.", "Segundo."]) # several texts
saber.extract([Path("a.txt"), Path("corpus/")]) # several paths
To read a plain string as a path, pass source_type="path".
In a folder, every plain-text file is read: files with a .txt extension and
files with no file-type extension at all. A trailing number is not treated
as an extension, so corpus names like A1.1 or texto.2 are read too. Hidden
files (.DS_Store, dotfiles) and binary files are skipped. To pick up something
else — or to narrow the selection — pass an explicit glob:
saber.extract(Path("corpus/"), pattern="*.txt") # only .txt
saber.extract(Path("corpus/"), pattern="*.tsv") # some other extension
saber.extract(Path("corpus/"), pattern="aluno_*") # a subset, by name
Other folder options: recursive=False, encoding="utf-8" (files that are not
valid UTF-8 fall back to latin-1 with a warning).
Each document is named after its file, without the extension (texto57.txt and
texto57 both become texto57, while A1.1 keeps its number); inline strings
become text, or text_1, text_2, … when there are several.
Choosing which structures to extract
Four independent filters. Values within one filter are OR-ed, and the filters are AND-ed together. Matching ignores case and diacritics.
# A broad group: by key or by its Portuguese name
saber.extract(text, groups="a5")
saber.extract(text, groups="Pronomes") # the same 29 structures
# A finer category
saber.extract(text, categories="a5d2") # Pronomes > Demonstrativos
# A CEFR level, or several
saber.extract(text, levels="B1")
saber.extract(text, levels=["A1", "A2"])
# Individual structures, by id or by matcher name
saber.extract(text, structures=["a5d2_1_A2", "b5d2_6"])
# Combined: subjunctive-level pronoun structures only
saber.extract(text, groups="a5", levels="B2")
# Everything except a few structures
saber.extract(text, exclude=["a1d1_7_A2"])
To see what the valid values are:
saber.list_structures() # all 252, as a DataFrame
saber.list_structures(groups="a7") # just the quantifiers
| structure | matcher | level | group | group_name | category | category_name | description |
|---|---|---|---|---|---|---|---|
| a5d2_1_A2 | a5d2_1 | A2 | a5 | Pronomes | a5d2 | Demonstrativos | Pronomes - Demonstrativos - contração com preposições |
An unknown filter value raises ValueError listing the valid options, and so
does a combination that selects nothing.
Structure ids encode the taxonomy, so you can also filter the output frame
directly: a5d2_1_A2 → group a5, category a5d2, matcher a5d2_1, level
A2.
Output columns
mode="spans" — one row per occurrence
| Column | Meaning |
|---|---|
document, path |
Which text the match came from (path is empty for inline strings) |
structure |
Full structure id, e.g. a5d2_1_A2 |
matcher |
Name of the matcher function, e.g. a5d2_1 |
level |
CEFR level, A1–C2 |
group, group_name |
Broad group key and Portuguese name |
category, category_name |
Category key and Portuguese name |
description |
Portuguese description of the structure |
text |
The matched substring, sliced out of the preprocessed text |
reconstructed_text |
The matcher's own rendering of the match |
token_start, token_end |
Token offsets, end-exclusive |
char_start, char_end |
Character offsets into the preprocessed text |
pipeline |
stanza or spacy, depending on which parse the matcher needs |
Rows are sorted by document and position. Structures overlap freely — a single phrase commonly matches several — so occurrences are not mutually exclusive.
text is the authoritative one: it is cut from the text by character offset, so
analysis.text[char_start:char_end] == text always holds and contractions come
out as written (ao, no, daquele). reconstructed_text comes from the
matcher's own reconstruct_text helper and is normally identical, but it is
rebuilt from tokens rather than sliced, so keep to text when exactness matters.
To get offsets against your own untouched string, note that they refer to the
preprocessed text — see How it works. saber.analyze()
returns that text alongside the matches if you need it:
analysis = saber.analyze("Ele foi ao cinema.", saber.all_structures())
analysis.text # the preprocessed text the offsets refer to
analysis.n_tokens # tokens excluding punctuation
analysis.matches # list of Match objects
mode="features" — one row per document
| Column | Meaning |
|---|---|
document, path |
Which text the row describes |
n_tokens |
Tokens excluding whitespace and punctuation |
<structure>_count |
Raw number of occurrences |
<structure>_norm |
Occurrences per 100 tokens |
<structure>_present |
1 if the structure occurs at all, else 0 |
Columns are emitted for every selected structure, including those that never
matched, so feature matrices line up across documents and across runs. Their
order follows the declaration order of LABEL_TRANSLATIONS in
saber/label_translation.py, which is curated taxonomically.
Selecting all 252 structures therefore gives 3 × 252 + 3 = 759 columns. Use
features= to keep only what you need:
saber.extract(corpus, mode="features", features=["normalized", "presence"])
How it works
- Preprocessing. Dialogue dashes are stripped and whitespace collapsed. All
offsets refer to this preprocessed text, which
saber.analyze()exposes asanalysis.text. - Parsing. The text is parsed twice: by
spaCy-Stanza (
tokenize,mwt,pos,lemma,depparse) for morphology and dependencies, and by spaCy'spt_core_news_mdfor the handful of matchers that need named entities or spaCy's own tokenisation. A matcher declares which parse it needs; you never choose. - Matching. Each selected structure's matcher runs over the appropriate
parse using spaCy's
Matcher,PhraseMatcherandDependencyMatcher, and returns token spans. - Offsets. Token spans are converted to character offsets. Stanza expands
multi-word tokens (
no→em o), which destroys spaCy's own character offsets, so they are recovered from the underlying Stanza document — this is whytextreadsnoand notem o. - Tabulating. Occurrences become the spans table, or are counted and normalised into the feature table.
A matcher that raises is reported as a warning and skipped, so one broken
pattern cannot abort a corpus run. Use on_error="raise" while debugging, or
on_error="ignore" to silence it.
Performance
Expect a few seconds per short text, and add a few additional seconds of one-off model loading
to the first extraction in a process. Two things dominate: the Stanza parse, and
the fact that each matcher rebuilds its spaCy Matcher object on every call.
Practical consequences:
- Extract once for all the structures you might want, then filter the resulting DataFrame — that is much cheaper than re-running with different filters.
- Restricting
groups/categories/levelsup front does cut the matching cost roughly in proportion, but not the parsing cost. - A progress bar over the documents appears whenever there is more than one
document and you are working in a terminal or a notebook.
progress=Falseturns it off,progress=Trueforces it on (useful when output is redirected). - Progress is also logged per document, which is the better option for scripts
and log files. Turn it on with
logging.basicConfig(level=logging.INFO).
Working on the matchers
A matcher function is named after its structure minus the CEFR suffix
(a5d2_1 → a5d2_1_A2), takes a spaCy Doc, and returns a list of
(label, text, token_start, token_end) tuples. Setting
<function>.REQUIRES_SPACY = True after the definition routes it to the
pt_core_news_md parse instead of the Stanza one.
A structure is available only if it has both an entry in
LABEL_TRANSLATIONS and a matcher function, so commenting out a label in
saber/label_translation.py disables the structure.
saber.missing_labels() lists labels with no matcher.
Testing the matchers
tests/tests.yaml holds positive and negative example sentences per matcher.
The harness measures precision and recall for each, both with all patterns
active and with only the first pattern registered:
pip install -e ".[dev]"
python tests/test_suite.py
Results are written to tests/test_results.csv; the committed copy is the
reference from the last full run.
Licence and attribution
Released under the MIT License — see LICENSE. You may use, modify, and distribute this work, including commercially, provided the copyright notice and licence text are retained.
The structure inventory and its CEFR mapping derive from the Referencial Camões PLE (Instituto Camões).
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file saber_pt-0.5.4.tar.gz.
File metadata
- Download URL: saber_pt-0.5.4.tar.gz
- Upload date:
- Size: 108.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c631e55eefbc56ab6cee1dd1ac9f93f5ea83173e6702b6af2266e3403b27139e
|
|
| MD5 |
2e8bb183246212162cd503b7d5abdde4
|
|
| BLAKE2b-256 |
ad511d546619a0c70742f7d84d0fb1edd4abe80f7658e5e8b7a295b9c73bb3fe
|
File details
Details for the file saber_pt-0.5.4-py3-none-any.whl.
File metadata
- Download URL: saber_pt-0.5.4-py3-none-any.whl
- Upload date:
- Size: 121.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
20775e0477dde3f177542e60f9456dfa70cfda46590ee627e51c229b38c61a92
|
|
| MD5 |
14d78a979d384793c82b12f64d5dec17
|
|
| BLAKE2b-256 |
267b0e51a21ce6e19d8b454e1af52f5ae33e0f8cea50107e312aeae3367b058d
|