Datalabs
Project description
DataLab API CN
Installation
Install
```shell
pip install --upgrade pip
pip install datalabs
```
or
```shell
pip install --upgrade pip
git clone https://github.com/ExpressAI/Datalab.git
cd Datalab
pip install .
```
Dataset Operation
# pip install datalab
from datalabs import operations, load_dataset
from featurize import *
dataset = load_dataset("ag_news")
# print(task schema)
print(dataset['test']._info.task_templates)
# data operators
res = dataset["test"].apply(get_text_length)
print(next(res))
# get entity
res = dataset["test"].apply(get_entity_spacy)
print(next(res))
# get postag
res = dataset["test"].apply(get_postag_spacy)
print(next(res))
from edit import *
# add typos
res = dataset["test"].apply(add_typo)
print(next(res))
# change person name
res = dataset["test"].apply(change_person_name)
print(next(res))
Task Schema
-
text-classification
text
:strlabel
:ClassLabel
-
text-matching
text1
:strtext2
:strlabel
:ClassLabel
-
summarization
text
:strsummary
:str
-
sequence-labeling
tokens
:List[str]tags
:List[ClassLabel]
-
question-answering-extractive
:context
:strquestion
:stranswers
:List[{"text":"","answer_start":""}]
one can use dataset[SPLIT]._info.task_templates
to get more useful task-dependent information, where
SPLIT
could be train
or validation
or test
.
Supported Datasets
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
datalabs-0.1.1.dev0.tar.gz
(318.4 kB
view hashes)
Built Distribution
Close
Hashes for datalabs-0.1.1.dev0-py2.py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 271221f9440074f1a445e3803bae2990f503baa66f263fbb6c2c1423c17a7d3d |
|
MD5 | fdf76b4988c8c5e998194c993b154fd4 |
|
BLAKE2b-256 | 497af2d72adb6fed197ab5facfdd54acec2a19f38aab38e1f44a9c7eacffa6d9 |