Skip to main content

Cover

BARRED: generate faithful and diverse training sets

PyPI Python versions License arXiv

Unofficial implementation of the BARRED paper:

Boundary Alignment Refinement through REflection and Debate, aka BARRED arXiv:2604.25203

[!NOTE] I developed this library for my own needs without endorsement. It is not affiliated with the authors of the paper.

What is BARRED?

In a nutshell: BARRED is a framework for generating faithful and diverse synthetic training data using only a task description and a small set of unlabeled examples.

Step by step

two_steps.png

  1. Define the task: a Criterion, plus a few examples of input you work on
    Criterion: True when the sentence expresses a positive sentiment, False otherwise
    
    Examples:
      - The delivery arrived two days late and the box was crushed.
      - Honestly, one of the best purchases I've made this year.
      - It works, I guess.
    
  2. Let the library decompose the problem into dimensions
  3. Then the library generates the training set

Results: a set of labeled samples:

1. The screws stripped on the first turn, total waste.  -> False
2. Arrived a day early, and better than the photos.     -> True
3. Does the job. Nothing more to say.                   -> False   <- the gray zone

Why is it cool?

Over the past few years, here's what kept happening to me:

  • I developed a classifier but got no dataset (business people were busy)
  • I asked for relevancy judgments on products, but got quite nothing
  • I begged for annotated samples of search queries but got none

So, BARRED is cool because it solves these problems by generating samples almost automatically.

And there is more. "Happy path" examples are easy to write. "Edge cases" are not, the gray zone, the "not true, not false".

Say you want to blacklist irrelevant products in search results (personal true story). Anyone can tell that a hammer drill is relevant for a query on drills. But what about drill bits? And a drill bit adapter? And a drill toy? And the dozen cases nobody thinks about?

That's where BARRED shines!

How?

The naive way is to let an LLM generate samples straight from the problem description. WRONG!

LLMs forget parts of the problem. They don't "explore" much, they stick to their few default answers.

First smart idea of BARRED

The first thing that BARRED does is to decompose the problem into "Dimensions" and "Instantiations" of those dimensions.

I didn't know what an "instantiation" was (except in computing), but an instantiation is "concrete evidence in support of a concept/claim".

Then, each instantiation will serve as a "seed" for the LLM to generate samples.

But can we generate samples directly from those instantiations right away? Of course not, that would be too simple!

Second smart idea of BARRED

generate_sample.png

BARRED makes two judges debate each sample until they agree. When a judge disagrees, the sample is reworked using their feedback, then debated again, and so on.

In the end, the generated samples can be accepted or not. This library only streams accepted samples, but you can also access the rejected ones using an Observer.

The Algorithm, complete edition

BARRED, as implemented:

1: Input: criterion C, examples E, target size N,
2:        concurrency K, max_debate_rounds B, max_refine_rounds R, max_attempts M = 2N
3:
4: # Step 1 — dimensions and their instantiations (paper lines 2-5, fused)
5: D ← decompose_dimensions(llm, criterion=C, examples=E, concurrency=5)
6:     # each d ∈ D is a DecomposedDimension(dimension, instantiations)
7:
8: # Step 2 — draw → generate → debate → refine, K attempts in flight
9: G ← ∅ ; attempts ← 0
10: while |G| < N and attempts < M do
11:     attempts ← attempts + 1
12:     d ~ Uniform(D)  ;  v ~ Uniform(d.instantiations)
13:     e ~ Uniform(E)  ;  y ~ Uniform({True, False})
14:
15:     s ← generate_sample(llm, criterion=C, example=e,
16:                         instantiation=v, target_verdict=y)
17:         # s is a Sample(reasoning, input_block, label)
18:
19:     for ℓ = 0, 1, ..., R do
20:         res ← debate(llm, criterion=C, sample=s, max_debate_rounds=B)
21:             # res is a DebateResult(valid, dissenting_feedback)
22:
23:         if res.valid then
24:             G ← G ∪ {s} ; yield s
25:             break
26:         end if
27:         if ℓ = R then break end if      # refining now would skip validation
28:
29:         s ← refine_sample(llm, criterion=C, example=e, instantiation=v,
30:                           sample=s, dissenting_feedback=res.dissenting_feedback)
31:     end for
32: end while
33: return G

The Algorithm, in bullets

(Because I like bullets, plus some notes)

  1. Task definition
    1. Criterion = task description = the classification Criterion.
      E.g., “Is the product relevant for the query?”
    2. Unlabeled examples
      10 to 30 are enough
      No label is needed, represent the “shape” of the input data.
      E.g., “query: drill, product: bosch hammer drill”
  2. Dimensions decomposition
    1. dimension extraction
    2. dimension deduplication // it may happen, at least the Authors added that step
    3. dimension instantiation
      For each dimension:
        instantiations = verbalized_sampling(dimension)
        // list of name + polarity + score
      
  3. Sample Generation
    1. Take a dimension
      Take an instantiation of this dimension
      Take an example
      Take a label (true/false)
      (not random, take each in turn, aka uniform distribution)
      Consequence:
      • binary label → 50/50 dataset → maybe not your reality
      • The “polarity” of the instantiation is discarded
    2. Generate a sample calling the LLM with these inputs
    3. The prompt enforces four simultaneous constraints:
      1. The sample must align with instantiation → diversity
      2. The label must align with the expected label → control
      3. The sample must match the example's domain and style → realism
      4. The sample must be a BOUNDARY CASE, not a trivial one.
        "stress-test a smart and successful classifier".
        • Trick: the generator also emits the reasoning = the justification of the label.
        • Trick: no meta-leakage allowed ("do not mention test cases, models, dimensions, or labels in your output")
        • Note: If the generate sample label is not the target one, the divergence is logged and ignored.
    4. Debate label
      • Two judges take the sample (text and label). One judge is precision-oriented, the other is recall-oriented. A judge provides a verdict (reasoning, label, confidence).
      • Round 1:
        • each judge classifies the text → get a judge label
        • Sample label = judge1 label = judge2 label → accept sample
      • Round 2 if not consensus:
        • Each judge receives its previous verdict + verdict of the other judge + reasoning of the sample (= why this sample should be like that)
        • Consensus? → accept sample or return the dissenting feedbacks
    5. Refine sample At this step, the sample was rejected and the judges provided feedback.
      A new sample is generated, and this sample is debated.
      = a sample is never accepted without a debate.
  4. Profit!

[!NOTE] What is Verbalized Sampling?

Do not ask for a list, ask for a “distribution”.
= instead of a list of strings (the instantiations), ask the LLM for a description of the instantiation, a polarity (true/false/both) and a score (probability). We don’t use the polarity, nor the score afterward.

Unlock Diversity and pushes beyond typical modes.

(This technique is at the core of BARRED and is awesome!)

Link to the Paper "Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity"

Installation

uv add barred

# with Pip
pip install barred

This library builds on any-llm from Mozilla AI, and ships no provider SDK by default.
You need the any-llm-sdk extra for the provider you want:

Provider Install
OpenAI uv add barred "any-llm-sdk[openai]"
Anthropic uv add barred "any-llm-sdk[anthropic]"
Gemini uv add barred "any-llm-sdk[gemini]"
All providers uv add barred "any-llm-sdk[all]"

See the full list of providers.

Then set the matching key in your environment (OPENAI_API_KEY, ANTHROPIC_API_KEY, ...) or pass it to the llm() constructor.

Show me some code

import asyncio

from barred import LLM, barred, decompose_dimensions

#
# Step 1. Task definition
#
criterion = "True when the sentence expresses a positive sentiment, False otherwise"
examples = [
    "The delivery arrived two days late and the box was crushed.",
    "Honestly, one of the best purchases I've made this year.",
    "It works, I guess.",
]


async def main():
    # Your provider, your model.
    # The key is read from the environment (OPENAI_API_KEY here), or with the ` api_key ` param.
    llm = LLM(provider="openai", model="gpt-5.6-terra")

    #
    # Step 2. Decompose the Criterion into Dimensions
    #
    dimensions = await decompose_dimensions(llm, criterion=criterion, examples=examples)

    #
    # Step 3. Generate Samples, streamed as they land
    #
    async for sample in barred(
        llm,
        criterion=criterion,
        examples=examples,
        dimensions=dimensions,
        num_samples=5,
    ):
        print(sample.label, "->", sample.input_block)


asyncio.run(main())

Notebook/Demo

[!NOTE] The notebook is the best way to see BARRED in action.

It includes: dimensions decomposition and the generation of samples.
You will also be able to give a grasp on the generated data

Open Notebook in GitHub Open Notebook in Google Colab

Authors and resources

Paper authors:

Company: Plurai.ai

The main article giving a high-level overview of BARRED (April 28, 2026): Introducing BARRED: turn any policy prompt into a high-accuracy efficient guardrail

[!NOTE] I highly recommend giving Plurai.ai a shot to create a classifier. Under the hood, the app uses (I think) a derived version of BARRED. The UI and the experience are great.
Try it out at plurai.ai.

Changelog/Releases

Changelog and releases are on GitHub Releases.

Development setup

Install after cloning the repository:

just install

Check everything (lint, type, tests...):

just checks

Release

Releases are triggered by a version tag: pushing v* runs .github/workflows/release.yml, which builds, smoke-tests the wheel and the sdist, then publishes to PyPI via Trusted Publishing (no token involved).

just release recipe does the whole sequence — bump, check, commit, tag, push:

just release minor   # 0.2.0 => 0.3.0
just release patch   # 0.2.0 => 0.2.1
just release rc      # 0.2.0 => 0.2.0rc1, published as a pre-release
just release 1.0.0   # explicit version

It refuses to start from a dirty tree or outside main, and rolls the bump back if anything fails, so a failed run leaves nothing behind.

Metadata

Release files for barred 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for barred 0.3.0
File Size Uploaded
barred-0.3.0.tar.gz 23.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for barred 0.3.0
File Interpreter ABI Platform
barred-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 51.1 kB

Release files / barred-0.3.0.tar.gz

Download URL barred-0.3.0.tar.gz
Size 23.2 kB
Tags Source
SHA-256 checksum
How to use checksums
d47520989fd71604ce3cfd9292471692f47200b32f506a64dd2dc5a8009054f8
BLAKE2b-256 checksum
How to use checksums
17e51bfe8dce055273cb8a31d6deb9fc8d99888f9d76c6c2d7837506a73fc98e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 8, 2026.

Transparency log

Release files / barred-0.3.0-py3-none-any.whl

Download URL barred-0.3.0-py3-none-any.whl
Size 27.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
583f6566e90839adbf47c0bf2687cf6a618f3a930a8afec69903a94bd5d4087f
BLAKE2b-256 checksum
How to use checksums
5eeea10d1c3d7d31023c966e44b9c143ef2a75ce03d0deee1c85dd0f85ea1768
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 8, 2026.

Transparency log

Release history Release notifications | RSS feed

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page