Skip to main content

langchain-doubleword

A LangChain integration package for Doubleword.

This package wires Doubleword's OpenAI-compatible inference API (https://api.doubleword.ai/v1) into LangChain and LangGraph as both real-time chat / embedding models and transparently-batched variants powered by autobatcher.

The batched variants are required to access models that Doubleword exposes only via the batch API, and they cut cost on workloads that fan out many concurrent calls, which is typical of LangGraph agents.

Installation

pip install langchain-doubleword

Authentication

Three resolution paths, in precedence order:

  1. Explicit constructor argument:

    ChatDoubleword(model="...", api_key="sk-...")
    
  2. Environment variable:

    export DOUBLEWORD_API_KEY=sk-...
    
  3. ~/.dw/credentials.toml: the same file written by Doubleword's CLI tooling. The active account is selected by ~/.dw/config.toml's active_account field, and inference_key from that account is used.

    # ~/.dw/config.toml
    active_account = "work"
    
    # ~/.dw/credentials.toml
    [accounts.work]
    inference_key = "sk-..."
    

    To use a non-active account from your credentials file, set DOUBLEWORD_API_KEY directly to that account's inference_key. There is no account= selector on the model itself.

Chat models

ChatDoubleword (real-time)

Drop-in chat model. Use this in any LangChain or LangGraph workflow that expects a BaseChatModel.

from langchain_doubleword import ChatDoubleword

llm = ChatDoubleword(model="your-model-name")

response = llm.invoke("Explain bismuth in three sentences.")
print(response.content)

ChatDoublewordBatch (transparently batched)

Same interface, but every concurrent .ainvoke() call is collected by autobatcher and submitted via Doubleword's batch endpoint. Async-only: sync .invoke() raises.

Use this when:

  • The model you want is batch-only (some Doubleword-hosted models do not expose a real-time chat endpoint).
  • You're running a LangGraph workflow with parallel branches and want ~50% cost savings via batch pricing.
import asyncio
from langchain_doubleword import ChatDoublewordBatch

llm = ChatDoublewordBatch(model="batch-only-model")

async def main():
    # Concurrent calls collected into a single batch under the hood.
    results = await asyncio.gather(*[
        llm.ainvoke(f"Summarize chapter {i}") for i in range(50)
    ])
    for r in results:
        print(r.content)

asyncio.run(main())

Tuning autobatcher

Four autobatcher.BatchOpenAI knobs are exposed as constructor arguments:

Argument Default Purpose
batch_size 1000 Submit a batch when this many requests are queued.
batch_window_seconds 10.0 Submit a batch after this many seconds even if the size cap is not reached.
poll_interval_seconds 5.0 How often autobatcher polls for batch completion.
completion_window "24h" Doubleword batch completion window. "1h" is more expensive but faster.
llm = ChatDoublewordBatch(
    model="your-model",
    batch_size=250,           # smaller batches for fast-turnaround LangGraph nodes
    batch_window_seconds=2.5, # don't make latency-sensitive calls wait 10s
    completion_window="1h",   # pay more, finish quicker
)

The same arguments are available on DoublewordEmbeddingsBatch.

ChatDoublewordAsync (1-hour flex tier)

A thin subclass of ChatDoublewordBatch pinned to Doubleword's flex (1-hour) completion window. Backed by autobatcher.AsyncOpenAI rather than BatchOpenAI. Use this when 24-hour batch turnaround is too slow but realtime cost is too high, which is typical for fan-out workflows that need results within minutes-to-an-hour.

import asyncio
from langchain_doubleword import ChatDoublewordAsync

llm = ChatDoublewordAsync(model="your-model")  # completion_window="1h" by default

async def main():
    results = await asyncio.gather(*[
        llm.ainvoke(f"Summarize chapter {i}") for i in range(50)
    ])
    for r in results:
        print(r.content)

asyncio.run(main())

All the autobatcher tuning knobs above apply unchanged. The only difference from ChatDoublewordBatch is the default completion_window ("1h" vs "24h"); the same DoublewordEmbeddingsAsync exists on the embeddings side.

Prompt caching

from langchain_doubleword import ChatDoubleword
from langchain_core.messages import SystemMessage, HumanMessage

llm = ChatDoubleword(model="your-model-name", cache_control={"type": "ephemeral", "ttl": "1h"})

response = llm.invoke([
    SystemMessage(content="…large, stable instructions…"),
    HumanMessage(content="What is 2 + 2?"),
])
print(response.usage_metadata["input_token_details"]["cache_read"])

ttl is "5m" or "1h" and is optional. The API default is "5m".

The marker goes on the last system message and on the latest message of every request. Hand-built cache_control blocks on system and user messages still work for any other placement.

Pass cache_control to invoke to override it for one call. cache_control=None skips caching for that call.

See the prompt caching guide.

Embeddings

from langchain_doubleword import (
    DoublewordEmbeddings,
    DoublewordEmbeddingsAsync,
    DoublewordEmbeddingsBatch,
)

embed = DoublewordEmbeddings(model="your-embedding-model")
vec = embed.embed_query("hello world")

# Or, transparently batched (24h tier):
batch_embed = DoublewordEmbeddingsBatch(model="your-embedding-model")
# vecs = await batch_embed.aembed_documents([...])

# Or on the 1h flex tier:
async_embed = DoublewordEmbeddingsAsync(model="your-embedding-model")
# vecs = await async_embed.aembed_documents([...])

Use with LangGraph

ChatDoubleword, ChatDoublewordBatch, and ChatDoublewordAsync are all standard BaseChatModel implementations, so they slot into any LangGraph node:

from langgraph.graph import StateGraph, END
from langchain_doubleword import ChatDoublewordBatch

llm = ChatDoublewordBatch(model="your-model")

async def call_model(state):
    response = await llm.ainvoke(state["messages"])
    return {"messages": [response]}

graph = StateGraph(dict)
graph.add_node("model", call_model)
graph.set_entry_point("model")
graph.add_edge("model", END)
app = graph.compile()

When several model nodes execute in parallel (e.g. via Send or fan-out edges), autobatcher collects their requests into a single batch.

Configuration

Argument Env var Default
api_key DOUBLEWORD_API_KEY required
base_url DOUBLEWORD_API_BASE https://api.doubleword.ai/v1
model n/a required

All other arguments accepted by langchain_openai.ChatOpenAI are forwarded unchanged (temperature, max_tokens, model_kwargs, timeout, etc.).

License

MIT

Metadata

Release files for langchain-doubleword 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for langchain-doubleword 0.3.0
File Size Uploaded
langchain_doubleword-0.3.0.tar.gz 15.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for langchain-doubleword 0.3.0
File Interpreter ABI Platform
langchain_doubleword-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 28.7 kB

Release files / langchain_doubleword-0.3.0.tar.gz

Download URL langchain_doubleword-0.3.0.tar.gz
Size 15.1 kB
Tags Source
SHA-256 checksum
How to use checksums
85df2a80b85eadf2923ae5d83b8b8537f937ce5fb716b8452fc7027fbfa665f5
BLAKE2b-256 checksum
How to use checksums
dbd9e2a16ef6ac7fa853ec9b2b29329ae2700ef865f4ec9be0a042e9203087e0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.

Transparency log

Release files / langchain_doubleword-0.3.0-py3-none-any.whl

Download URL langchain_doubleword-0.3.0-py3-none-any.whl
Size 13.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1f88b658c30e204fb8da796dc27470a20c4585525259a7f0db080dafb6d30499
BLAKE2b-256 checksum
How to use checksums
2383eb8fcb78e2b103f3d1457b6b99a59b8927abfc3d6a2f07badf0eaedcfdae
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.1

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page