Skip to main content

langchain-flexai

LangChain integration for FlexAI — open-weight models served behind an OpenAI-compatible API.

Installation

pip install -U langchain-flexai
export FLEXAI_API_KEY="your-api-key"

Chat models

from langchain_flexai import ChatFlexAI

llm = ChatFlexAI(model="DeepSeek-V4-Flash-0731")
llm.invoke("Explain speculative decoding in two sentences.")

model takes the canonical id as returned by GET /v1/models — the bare model name, without the organisation prefix.

Tool calling

from pydantic import BaseModel, Field


class GetWeather(BaseModel):
    """Get the current weather in a given location."""

    location: str = Field(description="City, e.g. Paris")


llm.bind_tools([GetWeather]).invoke("Weather in Paris?")

Structured output

llm.with_structured_output(GetWeather).invoke("Weather in Paris?")

Not every served model enforces a strict JSON schema. Models that do not will reject the request rather than silently ignore it — check a model's supported_parameters in GET /v1/models.

Reasoning models

Some FlexAI-served models return a reasoning trace in a reasoning_content field. That field is not part of the OpenAI schema, so ChatOpenAI discards it; this package surfaces it on both invoke and stream:

result = ChatFlexAI(model="gpt-oss-120b").invoke("What is 17*23?")
result.additional_kwargs["reasoning_content"]

Not every reasoning model uses the field — some reason inline in content instead — so treat it as present-or-absent rather than guaranteed.

Capabilities vary by model

FlexAI serves many models through one endpoint, so some behaviour is per-model rather than provider-wide. Verified against the live API:

Behaviour Notes
Streaming token usage Works. This package sets stream_usage=True by default, unlike ChatOpenAI, because FlexAI only reports usage when stream_options.include_usage is sent.
Forced tool choice Per-model, and best-effort rather than constrained decoding. DeepSeek-V4-Flash-0731 and gpt-oss-120b honour tool_choice="any" on a prompt that invites no tool call; gemma-4-31b-it declines, and the gateway returns 400 tool_choice_not_honored rather than forcing one.
Structured output Strict JSON schema is enforced on a subset of models. Those that do not support it reject the request rather than silently ignoring it.
Image input Supported on vision models only. A model that is not a vision model will not read the image.

Check a model's supported_parameters in GET /v1/models before relying on any of these.

Configuration

Variable Default Purpose
FLEXAI_API_KEY — API key. Required.
FLEXAI_API_BASE https://api.flex.ai/v1 Endpoint, for regional or self-hosted deployments.

Relationship to langchain-openai

FlexAI is OpenAI-compatible, so this package is a thin configuration of BaseChatOpenAI rather than a separate client. Everything ChatOpenAI supports works here. Using ChatOpenAI with base_url="https://api.flex.ai/v1" remains equivalent and supported; this package just supplies the defaults.

Documentation

docs.flex.ai/inference-api/agents/langchain

License

MIT

Metadata

Release files for langchain-flexai 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for langchain-flexai 0.1.0
File Size Uploaded
langchain_flexai-0.1.0.tar.gz 9.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for langchain-flexai 0.1.0
File Interpreter ABI Platform
langchain_flexai-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 17.2 kB

Release files / langchain_flexai-0.1.0.tar.gz

Download URL langchain_flexai-0.1.0.tar.gz
Size 9.7 kB
Tags Source
SHA-256 checksum
How to use checksums
4f5a62059ca7ed6779f9bf34277704e6242d6a7b7285528d1c4d5aa7e70751a4
BLAKE2b-256 checksum
How to use checksums
e8e68f53cf3f517fcb864a9a3a369fe773d3c67a6f1a155eb992d5653002f922
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / langchain_flexai-0.1.0-py3-none-any.whl

Download URL langchain_flexai-0.1.0-py3-none-any.whl
Size 7.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bda8635e728adbfdbe6e3c559cc98e2bba269483ab86af25e2f5ef039c5b6932
BLAKE2b-256 checksum
How to use checksums
ba3b9f658724bd5eba4b0e1de8ebdb2fe7888392b1323557fe423fcaa0ac0248
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page