Generation of Realistic Tabular data
with pretrained Transformer-based language models
Our GReaT framework leverages the power of advanced pretrained Transformer language models to produce high-quality synthetic tabular data. Generate new data samples effortlessly with our user-friendly API in just a few lines of code. Please see our publication for more details.
我们的GReaT框架利用先进的预训练Transformer语言模型的力量,生成高质量的合成表格数据。只需几行代码,就可以使用我们的用户友好的API轻松生成新的数据样本。更多详情请参阅我们的出版物
GReaT Installation
The GReaT framework can be easily installed using with pip - requires a Python version >= 3.9:
pip install be-great
GReaT Quickstart
In the example below, we show how the GReaT approach is used to generate synthetic tabular data for the California Housing dataset.
from be_great import GReaT
from sklearn.datasets import fetch_california_housing
data = fetch_california_housing(as_frame=True).frame
model = GReaT(llm='distilgpt2', batch_size=32, epochs=25)
model.fit(data)
synthetic_data = model.sample(n_samples=100)
Imputing a sample
GReaT also features an interface to impute, i.e., fill in, missing values in arbitrary combinations. This requires a trained model, for instance one obtained using the code snippet above, and a pd.DataFrame where missing values are set to NaN.
A minimal example is provided below:
# test_data: pd.DataFrame with samples from the distribution
# model: GReaT trained on the data distribution that should be imputed
# Drop values randomly from test_data
import numpy as np
for clm in test_data.columns:
test_data[clm]=test_data[clm].apply(lambda x: (x if np.random.rand() > 0.5 else np.nan))
imputed_data = model.impute(test_data, max_length=200)
GReaT Citation
If you use GReaT, please link or cite our work:
@inproceedings{borisov2023language,
title={Language Models are Realistic Tabular Data Generators},
author={Vadim Borisov and Kathrin Sessler and Tobias Leemann and Martin Pawelczyk and Gjergji Kasneci},
booktitle={The Eleventh International Conference on Learning Representations },
year={2023},
url={https://openreview.net/forum?id=cEygmQNOeI}
}
GReaT Acknowledgements
We sincerely thank the HuggingFace 🤗 framework.
Release files for be-great-v 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| be-great-v-0.1.3.tar.gz | 16.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| be_great_v-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 33.5 kB
Release files / be-great-v-0.1.3.tar.gz
| Download URL | be-great-v-0.1.3.tar.gz |
|---|---|
| Size | 16.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
af69f2422aae7a177198abb5d257ca0a1ab7e56f091c3a0368f35982f7170389
|
|
BLAKE2b-256 checksum How to use checksums |
48da6a7c19e1d1052349bcd739dacaf570335a50dcfc08b76ca34994f282a342
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.9.18
|
Release files / be_great_v-0.1.3-py3-none-any.whl
| Download URL | be_great_v-0.1.3-py3-none-any.whl |
|---|---|
| Size | 16.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9be8ca6259f72198bbb6f3f303d2a21a1cc428a184c0d40d83dfd8a4e6af8de9
|
|
BLAKE2b-256 checksum How to use checksums |
b8976274b762ebe90d6380129db875896ae120346d17f4b4735daaf5b356fce3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.9.18
|