Skip to main content

thirthfloor

Local AI inference for Python. One install. No daemon. No API keys.

Your model runs as an object inside your process. Ask a question, get a string back.

from thirthfloor import Engine

engine = Engine()
engine.load("qwen", "/models/qwen3-4b-q4.gguf")
print(engine.chat("qwen", "What is boundary value analysis?"))

Install

pip install 3thfloor

CPU works out of the box. For GPU acceleration, install the matching wheel for your hardware:

Apple Silicon (Metal)

pip install llama-cpp-python --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/metal
pip install 3thfloor

NVIDIA (CUDA 12.x)

pip install llama-cpp-python --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu124
pip install 3thfloor

Same package, same API. The backend picks up your GPU automatically.

Quick Start

from thirthfloor import Engine

engine = Engine()
engine.load("tester", "/models/3thfloor-tester-q4.gguf")

answer = engine.chat("tester", "Write three test cases for a login form.")
print(answer)

chat() returns a plain string. Not a dict, not .choices[0].message.content. A string.

Sessions (Conversation History)

session = engine.session("tester", system="You are a senior QA engineer.")

print(session.send("What is exploratory testing?"))
print(session.send("How is that different from what I asked about?"))

The session tracks history for you. The second question knows about the first. No message arrays to build, no history to splice.

Streaming

for token in engine.stream("tester", "Explain risk-based testing in two paragraphs."):
    print(token, end="", flush=True)

stream() yields token strings as they generate. Print them, pipe them, collect them.

Multiple Models

engine.load("fast", "/models/qwen3-4b-q4.gguf")
engine.load("smart", "/models/qwen3-32b-q4.gguf")

def ask(question: str) -> str:
    alias = "smart" if len(question) > 200 else "fast"
    return engine.chat(alias, question)

Load as many models as your RAM allows. Route between them with the alias.

Agents

from thirthfloor import tool, run_agent

@tool
def get_build_status(pipeline: str) -> str:
    """Return the latest CI status for a pipeline."""
    return "pipeline main: passing, 312 tests green"

result = run_agent(
    engine, "tester",
    "Is the main pipeline healthy?",
    tools=[get_build_status],
)
print(result)

Decorate a function with @tool, pass it to run_agent(). The model decides when to call it, the engine runs it, you get the final answer as a string. Docstrings and type hints become the tool schema.

Model Management

engine.models.add("tester", "/models/3thfloor-tester-q4.gguf")

for m in engine.models.list():
    print(m["alias"], m["path"], round(m["size_mb"] / 1024, 1), "GB")

engine.models.download("Qwen/Qwen3-4B-GGUF", filename="qwen3-4b-q4_k_m.gguf")

Registered models let you look up the path by alias: engine.load("tester", engine.models.info("tester")["path"]).

Downloading from HuggingFace requires the extra:

pip install "thirthfloor[hf]"

Optional HTTP Server

engine.serve(port=7437)

OpenAI-compatible endpoints (/v1/chat/completions) for when other tools need HTTP access. Never required. Requires:

pip install "thirthfloor[server]"

Embed in Your Software

from thirthfloor import Engine

class SupportBot:
    def __init__(self, model_path: str):
        self.engine = Engine()
        self.engine.load("support", model_path)
        self.session = self.engine.session(
            "support",
            system="You answer questions about our test automation product.",
        )

    def reply(self, message: str) -> str:
        return self.session.send(message)

bot = SupportBot("/models/3thfloor-tester-q4.gguf")
print(bot.reply("How do I tag a flaky test?"))

The model is an object in your process. No ports to manage, no subprocesses to babysit, no service that has to be running before your app starts. When your process exits, the model is gone. Ship it inside a CLI, a desktop app, a batch job, anywhere Python runs.


Built by Justin Bench, 3th Floor AI.

Free for personal projects, research, experiments, and noncommercial use under the PolyForm Noncommercial License 1.0.0. Commercial use requires a license: justin@3thfloor.com.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

3thfloor-0.1.1.tar.gz (13.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

3thfloor-0.1.1-py3-none-any.whl (16.0 kB view details)

Uploaded Python 3

File details

Details for the file 3thfloor-0.1.1.tar.gz.

File metadata

  • Download URL: 3thfloor-0.1.1.tar.gz
  • Upload date:
  • Size: 13.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.15

File hashes

Hashes for 3thfloor-0.1.1.tar.gz
Algorithm Hash digest
SHA256 afe6feec62bd59b924a7b9eca4f22c581c2f5557e5ae0734c243381bfbfc9a30
MD5 7d2bb7ce0df7c990c0029b95efe95a33
BLAKE2b-256 98a936a9e6a8c0e5043420c3485b56092ada86abfe9ef4a1309f9964ecfc0235

See more details on using hashes here.

File details

Details for the file 3thfloor-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: 3thfloor-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 16.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.15

File hashes

Hashes for 3thfloor-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 c4b3716352224cb8593f3bbfb9a3351955eb083bca9ab8b407550d29c1544a06
MD5 1e7922e6c5c1b73b83cf35f461b0b959
BLAKE2b-256 a78570e95e46e7ca2fd5e41c9a8228c8e647c88ad4ccf47a74da177d28f8c145

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page