A high-level NLP toolkit built on top of modern LLMs.
Project description
TextTools
📌 Overview
TextTools is a high-level NLP toolkit built on top of LLMs.
It provides three API styles for maximum flexibility:
- Sync API (
TheTool) - Simple, sequential operations - Async API (
AsyncTheTool) - High-performance async operations - Batch API (
BatchTheTool) - Process multiple texts in parallel with built-in concurrency control
It provides ready-to-use utilities for translation, question detection, categorization, NER extraction, and more - designed to help you integrate AI-powered text processing into your applications with minimal effort.
✨ Features
TextTools provides a collection of high-level NLP utilities. Each tool is designed to work with structured outputs.
categorize()- Classify text into given categoriesextract_keywords()- Extract keywords from the textextract_entities()- Perform Named Entity Recognition (NER)is_question()- Detect if the input is phrased as a questionto_question()- Generate questions from the given text / subjectmerge_questions()- Merge multiple questions into oneaugment()- Rewrite text in different augmentationssummarize()- Summarize the given texttranslate()- Translate text between languagespropositionize()- Convert a text into atomic, independent, meaningful sentencesis_fact()- Check whether a statement is a fact based on the source textrun_custom()- Custom tool that can do almost anything
🚀 Installation
Install the latest release via PyPI:
pip install -U hamtaa-texttools
📊 Tool Quality Tiers
| Status | Meaning | Tools | Safe for Production? |
|---|---|---|---|
| ✅ Production | Evaluated and tested. | categorize(), extract_keywords(), extract_entities(), is_question(), to_question(), merge_questions(), augment(), summarize(), run_custom() |
Yes - ready for reliable use. |
| 🧪 Experimental | Added to the package but not fully evaluated. | translate(), propositionize(), is_fact() |
Use with caution |
⚙️ Additional Parameters
-
with_analysis: bool→ Adds a reasoning step before generating the final output. Note: This doubles token usage per call. -
logprobs: bool→ Returns token-level probabilities for the generated output. You can also specifytop_logprobs=<N>to get the top N alternative tokens and their probabilities.
Note: This feature works if it's supported by the model. -
output_lang: str→ Forces the model to respond in a specific language. -
user_prompt: str→ Allows you to inject a custom instruction into the model alongside the main template. -
temperature: float→ Determines how creative the model should respond. Takes a float number between0.0and2.0. -
normalize: bool→ Whether to apply text cleaning (removing separator lines and normalizing quotation marks) before sending to the LLM. -
validator: Callable (Experimental)→ Forces the tool to validate the output result based on your validator function. Validator should return a boolean. If the validator fails, TheTool will retry to get another output by modifyingtemperature. You can also specifymax_validation_retries=<N>. -
priority: int (Experimental)→ Affects processing order in queues.
Note: This feature works if it's supported by the model and vLLM. -
timeout: float→ Maximum time in seconds to wait for the response before raising a timeout error.
Note: This feature is only available inAsyncTheTool. -
raise_on_error: bool→ (TheTool/AsyncTheTool) Raise errors (True) or return them in output (False). Default is True. -
max_concurrency: int→ (BatchTheToolonly) Maximum number of concurrent API calls. Default is 5.
🧩 ToolOutput
Every tool of TextTools returns a ToolOutput object which is a BaseModel with attributes:
-
result: Any -
analysis: str -
logprobs: list -
errors: list[str] -
ToolOutputMetadatatool_name: strprocessed_by: strprocessed_at: datetimeexecution_time: floattoken_usage: TokenUsagecompletion_usage: CompletionUsageprompt_tokens: intcompletion_tokens: inttotal_tokens: int
analyze_usage: AnalyzeUsageprompt_tokens: intcompletion_tokens: inttotal_tokens: int
total_tokens: int
-
Serialize output to JSON using the
model_dump_json()method. -
Verify operation success with the
is_successful()method. -
Convert output to a dictionary with the
model_dump()method.
Note: For BatchTheTool: Each method returns a list[ToolOutput] containing results for all input texts.
🧨 Sync vs Async vs Batch
| Tool | Style | Use Case | Best For |
|---|---|---|---|
TheTool |
Sync | Simple scripts, sequential workflows | • Quick prototyping • Simple scripts • Sequential processing • Debugging |
AsyncTheTool |
Async | High-throughput applications, APIs, concurrent tasks | • Web APIs • Concurrent operations • High-performance apps • Real-time processing |
BatchTheTool |
Batch | Process multiple texts efficiently with controlled concurrency | • Bulk processing • Large datasets • Parallel execution • Resource optimization |
⚡ Quick Start (Sync)
from openai import OpenAI
from texttools import TheTool
client = OpenAI(base_url="your_url", API_KEY="your_api_key")
model = "model_name"
the_tool = TheTool(client=client, model=model)
detection = the_tool.is_question("Is this project open source?")
print(detection.model_dump_json())
⚡ Quick Start (Async)
import asyncio
from openai import AsyncOpenAI
from texttools import AsyncTheTool
async def main():
async_client = AsyncOpenAI(base_url="your_url", api_key="your_api_key")
model = "model_name"
async_the_tool = AsyncTheTool(client=async_client, model=model)
translation_task = async_the_tool.translate("سلام، حالت چطوره؟", target_language="English")
keywords_task = async_the_tool.extract_keywords("This open source project is great for processing large datasets!")
(translation, keywords) = await asyncio.gather(translation_task, keywords_task)
print(translation.model_dump_json())
print(keywords.model_dump_json())
asyncio.run(main())
⚡ Quick Start (Batch)
import asyncio
from openai import AsyncOpenAI
from texttools import BatchTheTool
async def main():
async_client = AsyncOpenAI(base_url="your_url", api_key="your_api_key")
model = "model_name"
batch_the_tool = BatchTheTool(client=async_client, model=model, max_concurrency=3)
categories = await batch_tool.categorize(
texts=[
"Climate change impacts on agriculture",
"Artificial intelligence in healthcare",
"Economic effects of remote work",
"Advancements in quantum computing",
],
categories=["Science", "Technology", "Economics", "Environment"],
)
for i, result in enumerate(categories):
print(f"Text {i+1}: {result.result}")
asyncio.run(main())
✅ Use Cases
Use TextTools when you need to:
- 🔍 Classify large datasets quickly without model training
- 🧩 Integrate LLMs into production pipelines (structured outputs)
- 📊 Analyze large text collections using embeddings and categorization
📄 License
This project is licensed under the MIT License - see the LICENSE file for details.
🤝 Contributing
We welcome contributions from the community! - see the CONTRIBUTING file for details.
📚 Documentation
For detailed documentation, architecture overview, and implementation details, please visit the docs directory.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file hamtaa_texttools-2.6.0.tar.gz.
File metadata
- Download URL: hamtaa_texttools-2.6.0.tar.gz
- Upload date:
- Size: 30.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.28 {"installer":{"name":"uv","version":"0.9.28","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2021553c85a4269d35376c1b57e6abba5abddfa8ab227ea00575aa667ba2dc0f
|
|
| MD5 |
e100a4399e28bde07ee39afc7e81e38d
|
|
| BLAKE2b-256 |
2a5675411ee660a779a12cbb8471f8485034e47f20f96e590ec91cf3cfd05d04
|
File details
Details for the file hamtaa_texttools-2.6.0-py3-none-any.whl.
File metadata
- Download URL: hamtaa_texttools-2.6.0-py3-none-any.whl
- Upload date:
- Size: 39.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.28 {"installer":{"name":"uv","version":"0.9.28","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b6c9bfb6c260fc3f9718b78880e6ce52be3e5b2c927e4a370f02321fc2efe6fd
|
|
| MD5 |
cc6d8f16823c89b61dfa595a84c19c9a
|
|
| BLAKE2b-256 |
067a1940eb055dd8ba2b1eaba81db19d3cd87990c082c9409d2eca38aae860a2
|