Skip to main content

A package to process CSV text data in batches using OpenAI API

Project description

LLM Batch Processor 🚀

A Python package to process CSV text data in batches using OpenAI's GPT API.

Python License OpenAI API


📌 Features

✅ Processes large CSV files (50K+ rows) in batches of 10 rows
✅ Uses OpenAI GPT-4 API (for now) to generate responses for each row
✅ Saves the results in a new CSV column
✅ CLI support (llm-process) for easy execution
✅ Modular and scalable package design


📂 Project Structure

llm_batch_processor/
│── llm_batch_processor/  # The package
│   │── __init__.py
│   │── config.py         # Configuration settings
│   │── api_client.py     # Handles OpenAI API requests
│   │── processor.py      # Batch processing logic
│── scripts/
│   │── main.py           # CLI entry point
│── data/                 # Place CSV files here
│── setup.py              # Package setup
│── setup.cfg
│── pyproject.toml
│── requirements.txt
│── README.md
│── .gitignore
│── LICENSE

🛠 Installation with repo

1️⃣ Clone the repository

git clone https://github.com/karthikRavichandran/llm_batch_processor.git
cd llm_batch_processor

Or

2️⃣ Install the package

pip install llm-batch-process

3️⃣ Set up OpenAI API Key

Create a .env file in the root directory and add:

OPENAI_API_KEY=your_api_key_here
Project_Id=your_project_id (optional)

or

export OPENAI_API_KEY="your_api_key_here"
export Project_Id="your_project_id" #Optional 

🚀 Usage

1️⃣ Prepare your CSV file

  • Place your input CSV file in the data/ folder.
  • Ensure it has a column containing text (default: "text").

2️⃣ Run the batch processor

llm-process

This will:

  • Read data from data/input.csv
  • Process text in batches of 10 rows
  • Send each batch to OpenAI GPT-4
  • Save responses in data/output.csv under a new column "response"

⚙️ Configuration

Modify config.py for:

BATCH_SIZE = 10  # Number of rows per batch
INPUT_CSV = "data/input.csv"
OUTPUT_CSV = "data/output.csv"
TEXT_COLUMN = "text"
RESPONSE_COLUMN = "response"

🛠 Development

Install dependencies

pip install -r requirements.txt

Run tests

(TODO: Add unit tests)

pytest

📦 Publish to PyPI

1️⃣ Build the package

python setup.py sdist bdist_wheel

2️⃣ Upload to PyPI

twine upload dist/*

3️⃣ Install from PyPI

pip install llm_batch_processor

📜 License

This project is licensed under the MIT License.


💡 Contributing

  1. Fork the repo
  2. Create a feature branch (git checkout -b feature-branch)
  3. Commit changes (git commit -m "Added new feature")
  4. Push to branch (git push origin feature-branch)
  5. Create a Pull Request 🚀

📧 Contact

For support or collaboration, reach out via Karthik Ravichandran.


Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llm_batch_processor-0.1.1.tar.gz (5.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

llm_batch_processor-0.1.1-py3-none-any.whl (7.3 kB view details)

Uploaded Python 3

File details

Details for the file llm_batch_processor-0.1.1.tar.gz.

File metadata

  • Download URL: llm_batch_processor-0.1.1.tar.gz
  • Upload date:
  • Size: 5.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.9.21

File hashes

Hashes for llm_batch_processor-0.1.1.tar.gz
Algorithm Hash digest
SHA256 a98f4f5870db3127893bfc8cc7ee26f338c644180396f173327d12cd64e146de
MD5 aa118d07251495816ca6d10ca18dc64e
BLAKE2b-256 012c4d9d2c0ec502439b66632e7f74efcfc55079064863ac7c8cabde0b9bf7d9

See more details on using hashes here.

File details

Details for the file llm_batch_processor-0.1.1-py3-none-any.whl.

File metadata

File hashes

Hashes for llm_batch_processor-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 607209e7d9246132cd21b6f8cbde1f15ae3129dc8e26f517b663e1a9d99a5805
MD5 fd88c16546db970e7200fa22f1967cbd
BLAKE2b-256 919793f9de5d6d0668a08e71a2f22d5230827877bc33f0f6695bd32c70e14a31

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page