Skip to main content

A CLASSLA Fork of Stanza For Processing South Slavic Languages

Installation

pip

We recommend that you install Classla via pip, the Python package manager. To install, run:

pip install classla

This will also help to resolve all dependencies.

Running Classla

Getting started

To run your first Classla pipeline, follow these steps:

>>> import classla
>>> classla.download('sl')                            # to download models in Slovene
>>> nlp = classla.Pipeline('sl')                      # to initialize default Slovene pipeline
>>> doc = nlp("France Prešeren je rojen v Vrbi.")     # to run pipeline
>>> print(doc.conll_file.conll_as_string())           # to print output in conllu format
# newpar id = 1
# sent_id = 1.1
# text = France Prešeren je rojen v Vrbi.
1	France	France	PROPN	Npmsn	Case=Nom|Gender=Masc|Number=Sing	4	nsubj	_	NER=B-per
2	Prešeren	Prešeren	PROPN	Npmsn	Case=Nom|Gender=Masc|Number=Sing	1	flat_name	_	NER=I-per
3	je	biti	AUX	Va-r3s-n	Mood=Ind|Number=Sing|Person=3|Polarity=Pos|Tense=Pres|VerbForm=Fin	4	cop	_	NER=O
4	rojen	rojen	ADJ	Appmsnn	Case=Nom|Definite=Ind|Degree=Pos|Gender=Masc|Number=Sing|VerbForm=Part	0	root	_	NER=O
5	v	v	ADP	Sl	Case=Loc	6	case	_	NER=O
6	Vrbi	Vrba	PROPN	Npfsl	Case=Loc|Gender=Fem|Number=Sing	4	obl	_	NER=B-loc|SpaceAfter=No
7	.	.	PUNCT	Z	_	4	punct	_	NER=O

You can also look into pipeline_demo.py file for usage examples.

Processors

Classla pipeline is built from multiple units. These units are called processors. By default classla runs tokenize, ner, pos, lemma and depparse processors.

You can specify which processors classla runs, with processors attribute as in the following example.

>>> nlp = classla.Pipeline('sl', processors='tokenize,ner,pos,lemma')

Tokenization (tokenize)

In case you already have tokenized text, you should split the text (with i.e. spaces) and pass attribute tokenize_pretokenized=True.

By default classla uses a rule-based tokenizer - reldi-tokeniser.

Most important attributes:

tokenize_pretokenized   - [boolean]     ignores tokenizer

Part-of-speech tagging (pos)

Pos tagging processor will create output, that will contain part-of-speech tags and other features presented on universal dependencies webiste . It is optional and requires you to use tokenize processor beforehand.

Lemmatisation (lemma)

Lemmatization processor will produce lemmas for each word in input. It requires the usage of both tokenize and pos processors.

Parsing (depparse)

Parsing processor (named depparse in code) creates connections between words explained on universal dependencies website . It requires tokenizer and pos processors.

NER (ner)

Ner processor will try to find named entities in text. It requires tokenize processor.

Release files for classla 0.0.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for classla 0.0.3
File Size Uploaded
classla-0.0.3.tar.gz 154.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for classla 0.0.3
File Interpreter ABI Platform
classla-0.0.3-py3-none-any.whl Python 3 none any Details

Total release size: 359.3 kB

Release files / classla-0.0.3.tar.gz

Download URL classla-0.0.3.tar.gz
Size 154.4 kB
Tags Source
SHA-256 checksum
How to use checksums
f6adb734a1384288421c048a8fac77937d919833ab7694df7f7568dbb08ff56a
BLAKE2b-256 checksum
How to use checksums
6cd123451d9374e9cc04d72b2df793dd65aabe63c2b6ed8e55fe88467c482b61
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.1.1 pkginfo/1.5.0.1 requests/2.21.0 setuptools/46.4.0 requests-toolbelt/0.9.1 tqdm/4.46.1 CPython/3.7.1

Release files / classla-0.0.3-py3-none-any.whl

Download URL classla-0.0.3-py3-none-any.whl
Size 204.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a370fdb5e471a1db4bb7d6f1f2b3b88ec9f363893c367d45103d071627849c43
BLAKE2b-256 checksum
How to use checksums
e84138334704e6696cf5d46c242ea45ef32bb219daa59e62baf44965bcc51a2e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.1.1 pkginfo/1.5.0.1 requests/2.21.0 setuptools/46.4.0 requests-toolbelt/0.9.1 tqdm/4.46.1 CPython/3.7.1

Release history Release notifications | RSS feed

2.2.3

2 release files

2.2.2

2 release files

2.2.1

2 release files

2.2

2 release files

2.1.1

2 release files

2.1

2 release files

2.0

2 release files

1.2.0

3 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.4

2 release files

This release

0.0.3 This release

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page