Skip to main content

Lightweight Python framework for building LLM agents with tool calling and RAG

Project description

llm-agent-base

A lightweight Python library for building LLM agents with tool calling and retrieval-augmented generation (RAG). Works with any OpenAI-compatible API (OpenRouter, OpenAI, Ollama, etc.).

Available on PyPI: https://pypi.org/project/llm-agent-base/

Installation

pip install llm-agent-base

Features

  • Simple LLM calls — single-call ask() with optional knowledge retrieval for straightforward completions
  • Agentic looprun() with knowledge retrieval and tool calling until a final text response
  • Conversational chatchat() maintains conversation history across calls; reset_conversation() starts fresh
  • File attachments — pass images, PDFs, and text files directly to ask(), run(), or chat() via the files parameter
  • Tool calling — register plain Python functions as LLM-callable tools; schemas are built automatically from type hints and docstrings
  • RAG — ingest a folder of documents (.txt, .md, .json, .pdf) into a FAISS vector index and inject relevant chunks into every prompt
  • Knowledge search tool — when a knowledge base is configured, the agent automatically gains a search_knowledge tool (semantic/vector) and a read_knowledge_files tool (keyword search returning full file contents); both are optional and independently toggleable
  • Pipelines — chain multiple agents so each agent's output becomes the next agent's input
  • Temperature control — set per-agent temperature for precise or creative responses
  • Response format — enforce JSON output or structured schemas via the OpenAI response format API
  • Debug mode — optional logging of tool calls and knowledge retrievals

Quick start

from llm_agent_base import AgentBase, LLMConnectionConfig

config = LLMConnectionConfig(model="openai/gpt-4o-mini", api_key="...")

agent = AgentBase(
    system_prompt="You are a helpful assistant.",
    llm_config=config,
)

# Simple one-shot call
print(agent.ask("What is the capital of France?"))

# Full agentic loop (knowledge retrieval + tool calling)
print(agent.run("What is the capital of France?"))

The default base_url points to OpenRouter, which gives access to many models through a single API key. You can swap it for the OpenAI base URL or any other compatible endpoint.

Usage

Simple LLM call

ask() makes a single completion call with no tool calling. If a knowledge base is configured, relevant chunks are automatically retrieved and injected as context.

from llm_agent_base import AgentBase, LLMConnectionConfig

config = LLMConnectionConfig(model="openai/gpt-4o-mini", api_key="...")

agent = AgentBase(system_prompt="You are a helpful assistant.", llm_config=config)
print(agent.ask("Summarise the water cycle in one sentence."))

Temperature and response format

# Low temperature for deterministic, factual responses
precise = AgentBase(
    system_prompt="You are a helpful assistant.",
    llm_config=config,
    temperature=0.1,
)

# High temperature for creative responses
creative = AgentBase(
    system_prompt="You are a poet.",
    llm_config=config,
    temperature=1.4,
)

# Enforce JSON output
json_agent = AgentBase(
    system_prompt="Always respond with valid JSON.",
    llm_config=config,
    response_format={"type": "json_object"},
)
print(json_agent.ask('Return {"city": "Paris", "country": "France"}'))

Tool calling

Register any Python function as a tool. The function name becomes the tool name, the docstring becomes its description, and the type hints define the parameter schema.

from llm_agent_base import AgentBase, LLMConnectionConfig

config = LLMConnectionConfig(model="openai/gpt-4o-mini", api_key="...")

agent = AgentBase(
    system_prompt="You are a helpful assistant. Use the available tools when needed.",
    llm_config=config,
)

@agent.register_tool
def get_weather(city: str) -> str:
    """Return the current weather for a given city."""
    return f"The weather in {city} is sunny and 22°C."

@agent.register_tool
def add(a: int, b: int) -> int:
    """Add two integers and return the result."""
    return a + b

print(agent.run("What is the weather in Tokyo and what is 10 + 20?"))

register_tool can also be called directly:

agent.register_tool(get_weather)

Conversational chat

chat() works like run() — knowledge retrieval and tool calling are both active — but it accumulates the conversation history across calls so the LLM retains context between turns. Call reset_conversation() to wipe the history and start a new session.

from llm_agent_base import AgentBase, LLMConnectionConfig

config = LLMConnectionConfig(model="openai/gpt-4o-mini", api_key="...")

agent = AgentBase(
    system_prompt="You are a helpful assistant. Keep answers concise.",
    llm_config=config,
)

print(agent.chat("My name is Alice."))
print(agent.chat("What is my name?"))  # agent remembers: Alice

agent.reset_conversation()

print(agent.chat("What is my name?"))  # history cleared, agent no longer knows

RAG (knowledge base)

Place your documents in a folder (organised into subdirectories by topic). Call ingest_knowledge once to embed and index them, then use run, ask, or chat — relevant chunks are automatically retrieved and injected into the system prompt on every call.

When a knowledge base is configured, the agent registers two tools the LLM can call mid-reasoning:

  • search_knowledge — semantic vector search; returns the most relevant chunks for a query. Requires an ingested FAISS index.
  • read_knowledge_files — keyword search by filename or file content; returns complete file text. Works directly on files without a vector index.

Both tools are enabled by default. Use knowledge_search_tool=False or knowledge_file_tool=False to disable either one.

knowledge/
├── products/
│   ├── faq.md
│   └── pricing.json
└── support/
    └── sla.md
from llm_agent_base import AgentBase, LLMConnectionConfig

config = LLMConnectionConfig(model="openai/gpt-4o-mini", api_key="...")

agent = AgentBase(
    system_prompt="You are a product assistant. Answer using only the provided context.",
    llm_config=config,
    knowledge_folder_path="knowledge",
    knowledge_index_dir=".kb_index",  # where the FAISS index is saved
    knowledge_top_k=3,                # number of chunks injected per prompt
)

# Load from disk if a saved index exists, otherwise ingest and save automatically
agent.load_or_ingest_knowledge()

print(agent.run("Who founded the company and when?"))

Pass auto_load_or_ingest=True to do this in the constructor:

agent = AgentBase(
    system_prompt="You are a product assistant. Answer using only the provided context.",
    llm_config=config,
    knowledge_folder_path="knowledge",
    auto_load_or_ingest=True,
)

print(agent.run("Who founded the company and when?"))

For more control, call ingest_knowledge() and load_knowledge() directly:

# Build and persist the index (run once, or when documents change)
agent.ingest_knowledge(save=True)

# On subsequent runs, load from disk instead of re-embedding
agent.load_knowledge()

Keyword file search without a vector index

read_knowledge_files searches filenames and file contents directly — no embedding or FAISS index needed. Disable search_knowledge to use only this tool:

agent = AgentBase(
    system_prompt="You are a product assistant. Answer using only the provided context.",
    llm_config=config,
    knowledge_folder_path="knowledge",
    knowledge_search_tool=False,  # no vector index required
)

print(agent.run("What are the support SLA terms?"))

The LLM can search with fine-grained control over where to look and how many keywords must match:

# filename or content, at least 2 of the 3 keywords must match
read_knowledge_files(
    keywords=["pricing plan", "enterprise", "discount"],
    search_in="both",
    match_mode="min",
    min_matches=2,
)
search_in Where keywords are matched
"filename" File name only
"content" File contents only
"both" (default) Filename first, then contents
match_mode Files returned
"any" (default) At least one keyword matches
"all" Every keyword must match
"min" At least min_matches keywords match

File attachments

All three call methods accept an optional files parameter — a list of file paths to attach to the prompt. Files are embedded directly in the message sent to the model.

Type Extensions
Images .png, .jpg, .jpeg, .gif, .webp
Documents .pdf
Text .txt, .md, .json, .csv, .xml, .html, .yaml, .yml, .py, .js, .ts

Passing a file with any other extension raises a ValueError.

# Describe an image (requires a vision-capable model)
print(agent.ask("What's in this diagram?", files=["architecture.png"]))

# Summarise a PDF
print(agent.run("Summarise the key findings.", files=["report.pdf"]))

# Multi-turn with an attached file — the file stays in conversation history
agent.chat("Here's our codebase overview.", files=["overview.md"])
agent.chat("Which module handles authentication?")

Agent pipelines

Chain agents so the output of one becomes the input of the next:

from llm_agent_base import AgentBase, AgentPipelineBase, LLMConnectionConfig

config = LLMConnectionConfig(model="openai/gpt-4o-mini", api_key="...")

researcher = AgentBase(
    system_prompt="Extract the key facts from the user's question.",
    llm_config=config,
)
writer = AgentBase(
    system_prompt="Turn the provided facts into a concise, friendly summary.",
    llm_config=config,
)

pipeline = AgentPipelineBase(agents=[researcher, writer])
print(pipeline.run("Tell me about the Acme Corp product lineup."))

Debug mode

Pass debug=True to any agent to print tool invocations and knowledge retrievals to stdout:

agent = AgentBase(..., debug=True)
[debug] Retrieving knowledge
[debug] tool 'get_weather' args={'city': 'Tokyo'} result=The weather in Tokyo is sunny and 22°C.

API reference

LLMConnectionConfig

Parameter Type Default Description
model str Model identifier (e.g. "openai/gpt-4o-mini")
base_url str OpenRouter API base URL
api_key str | None None API key (falls back to OPENROUTER_API_KEY env var)
embedding_model str "openai/text-embedding-3-small" Model used for RAG embeddings

AgentBase

Parameter Type Default Description
system_prompt str System prompt sent on every call
llm_config LLMConnectionConfig Connection and model settings
temperature float | None None Sampling temperature (model default when omitted)
response_format dict | None None OpenAI response format (e.g. {"type": "json_object"})
knowledge_folder_path str | None None Folder of documents to index for RAG
knowledge_index_dir str ".kb_index" Directory where the FAISS index is persisted
knowledge_top_k int 5 Number of chunks injected per prompt
auto_load_or_ingest bool False Load saved index on init, or ingest and save if none exists
knowledge_search_tool bool True Register the search_knowledge semantic vector search tool
knowledge_file_tool bool True Register the read_knowledge_files keyword file search tool
debug bool False Print tool calls and retrievals to stdout
Method Description
ask(prompt, files) Single LLM call with optional knowledge retrieval; no tool calling
run(prompt, files) Full agentic loop — knowledge retrieval + tool calling until text response
chat(message, files) Like run() but accumulates conversation history across calls
reset_conversation() Clear the stored conversation history
register_tool(fn) Register a function as a tool; usable as a decorator
ingest_knowledge(save) Parse, embed, and index documents in knowledge_folder_path
load_knowledge() Restore a previously saved index from knowledge_index_dir
load_or_ingest_knowledge() Load saved index if one exists, otherwise ingest and save
retrieve_knowledge(query) Manually retrieve the top-k chunks for a query

read_knowledge_files tool parameters (called by the LLM, not directly):

Parameter Type Default Description
keywords list[str] Phrases to search for; each entry matched as-is
search_in str "both" "filename", "content", or "both"
match_mode str "any" "any" (OR), "all" (AND), or "min" (at least min_matches)
min_matches int | None None Minimum number of matching keywords when match_mode="min"

Other exports

Class / function Description
AgentPipelineBase Chains multiple AgentBase instances in sequence
KnowledgeBase Document ingestion, embedding, and FAISS retrieval
DocumentChunk Dataclass representing a retrieved text chunk
build_tool_schema Builds an OpenAI-compatible tool schema from a function
execute_tool_loop Runs the agentic tool-calling loop against any OpenAI-compatible client

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llm_agent_base-0.2.3.tar.gz (15.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

llm_agent_base-0.2.3-py3-none-any.whl (13.3 kB view details)

Uploaded Python 3

File details

Details for the file llm_agent_base-0.2.3.tar.gz.

File metadata

  • Download URL: llm_agent_base-0.2.3.tar.gz
  • Upload date:
  • Size: 15.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for llm_agent_base-0.2.3.tar.gz
Algorithm Hash digest
SHA256 deef4f8aa4da3da1f8e6b5c193b9323325339de5c6a2aaecf8d30016ffc1d7bb
MD5 4112cc1c3719a0f236ce2ad5a196074e
BLAKE2b-256 a177bb931728b0222ea2c44742f5c67b848bdb1368df68d2b8f7d0a757713f77

See more details on using hashes here.

Provenance

The following attestation bundles were made for llm_agent_base-0.2.3.tar.gz:

Publisher: publish.yml on mksochota16/llm-agent-base

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file llm_agent_base-0.2.3-py3-none-any.whl.

File metadata

  • Download URL: llm_agent_base-0.2.3-py3-none-any.whl
  • Upload date:
  • Size: 13.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for llm_agent_base-0.2.3-py3-none-any.whl
Algorithm Hash digest
SHA256 e28f0e587a36844d1bf75d96bb9f5e0581d5d06a69a880385be77aee0307cae3
MD5 2a644b104ba96f02d08c022cffef4fd0
BLAKE2b-256 9c52eab056223c22ccc9e39a422b2d6411cb25eab91bc76d821f80d477b77b36

See more details on using hashes here.

Provenance

The following attestation bundles were made for llm_agent_base-0.2.3-py3-none-any.whl:

Publisher: publish.yml on mksochota16/llm-agent-base

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page