Foclan 1.0: a compact LLM-first language for precise data transformation and output shaping.
Project description
Foclan 1.0
Foclan is a compact LLM-first orchestration language for precise data transformation, exact JSON-like output shaping, and agent-friendly workflow steps.
The core idea is simple: in vibecoding, not only the prompt and the model matter. The language itself is also a variable.
Most mainstream languages were optimized for human writing and reading. That does not automatically make them optimal for LLMs. Foclan explores the opposite direction: a language designed around what current LLMs tend to do well, while removing some of the things that are natural for humans but fragile for models.
This repository is the standalone public Foclan 1.0 package. It contains:
- the installable
foclanCLI - the stable recommended dialect
- the packaged prompt bundle
- executable examples
- Codex and Cursor integration scaffolds
The Story
Most people working with vibecoding currently optimize two things:
- the prompt
- the model
Foclan is built around a third variable:
- the language the model is asked to think in
That is the core bet behind the project.
Human-oriented languages like Python are excellent for human authors. But an LLM is not a human author. It is a token predictor with strong priors, limited working memory, and a tendency to overbuild code when the language gives it too many degrees of freedom.
Foclan tries to answer a narrow but important question:
What if the language itself was designed to make LLM-generated programs shorter, more exact, and less brittle?
That is why Foclan is intentionally biased toward:
- linear start-to-end flow
- exact output shaping
- very small state surface
- fewer moving pieces per program
- compact solutions that are easy to validate
Why Foclan Exists
Foclan is based on the hypothesis that LLM code quality can improve not only through:
- a better prompt
- a better model
but also through:
- a better language for the model
The current focus is not on replacing a general-purpose language. The focus is on giving LLMs a cleaner medium for:
- simple but exact data manipulation
- filtering, grouping, counting, sorting, and selection
- building precise JSON-like outputs
- preparing compact payloads for downstream API or LLM calls
- orchestrating narrow agent steps in a predictable way
The updated product direction is:
- keep the core language extremely small and LLM-optimized
- add practical capability through extensions
- and eventually allow controlled bridging into other languages when Foclan is not the right expression tool for a step
Why A Developer Might Actually Care
Foclan is not primarily a language for people to write by hand.
It is a language for LLMs to write inside real workflows.
The practical developer story is:
- keep Python, TypeScript, SQL, or your normal stack for the main application
- let the LLM use Foclan for the narrow parts where exact shape and low failure surface matter most
In practice that means:
- the human does not need to adopt Foclan as a new general-purpose language
- the human gives the model a smaller, more disciplined medium for exact transformation steps
- the surrounding product or application can still stay in the existing stack
So the point is not human ergonomics first.
The point is to give the model a better language for:
- exact data workflow
- LLM-first glue code
- agent pipeline steps
- compact transformation layers between systems
Why Foclan Can Be Easier For LLMs
Foclan tries to reduce some of the things that often make LLM-generated programs brittle:
- linear start-to-end flow
- one current value,
focus, instead of many mutable variables - minimal state tracking across distant parts of the program
- no traditional loops or recursion in the recommended style
- exact output shaping as a first-class concern
- compact syntax for the most common data tasks
The intended effect is:
- fewer syntax failures
- fewer logic mistakes in output shape
- shorter generated programs
- lower output token counts
- lower latency when the model uses the language well
Best Use-Cases
The best current use-cases are the ones where Python is expressive for humans but too open-ended for LLMs.
Good current fits:
- exact JSON response shaping
- nested report building from multiple inputs
- filtering, counting, grouping, sorting, and top-selection
- provider payload assembly before downstream API calls
- schema-driven extraction pipelines
- deterministic glue code between raw inputs and structured outputs
- small "dashboard" style programs where key names and nesting must be exact
- LLM extraction followed by deterministic cleanup and reshaping
- compact orchestration steps inside larger agent workflows
- data pipelines where the output contract matters more than general-purpose expressiveness
Typical places where LLMs often do worse in Python than in Foclan:
- they add extra wrapper objects like
report,summary, orresult - they return the right data under the wrong key names
- they overuse helper variables and drift away from the requested output shape
- they overengineer simple transforms into longer code with more failure surface
- they make local logic mistakes while juggling state across multiple intermediate variables
- they write plausibly correct code that is still structurally wrong for downstream systems
Foclan is especially promising when the real target is not "general coding" but:
- "return exactly this object"
- "compose these few inputs into this exact contract"
- "make the LLM stop improvising structure"
In other words, Foclan is best thought of as:
- a language for linearly transforming data
- a language for shaping exact outputs
- a language for orchestrating narrow steps in LLM-heavy workflows
Benchmark Signals So Far
Foclan is still early, but the benchmark results are already interesting.
Selected internal results:
| Setup | Foclan | Python |
|---|---|---|
| GPT-5-mini, minimal, main suite | 72.33% accuracy, 140.9 output tokens, 2.27s | 66.35% accuracy, 192.0 output tokens, 2.76s |
| GPT-5-mini, minimal, blind holdout | 68.14% accuracy, 158.4 output tokens, 2.42s | 59.80% accuracy, 211.6 output tokens, 2.89s |
GPT-5.4 exploratory heavy sample, Foclan none vs Python low |
30.0% accuracy, 422.6 visible code tokens, 426.6 billed output tokens, 5.05s | 16.67% accuracy, 495.2 visible code tokens, 661.5 billed output tokens, 8.69s |
These are not universal claims. They are methodology-specific internal benchmarks. But they do suggest that for a meaningful class of hard transformation tasks, Foclan can already outperform Python on:
- accuracy
- visible code size
- billed output size
- end-to-end latency
Important Tradeoffs
Foclan is not intended to be the nicest language for humans to write or read directly.
It is also not optimized for runtime execution efficiency in the way a mature general-purpose language is. The design target is LLM generation quality first, not raw execution speed of the produced programs.
There are also real current drawbacks:
- LLMs still have to learn Foclan from the prompt each time
- reasoning-capable models often overthink Foclan
- that overthinking can spend many hidden reasoning tokens
- Python still has a huge familiarity advantage from training data
Foclan also depends heavily on having good extensions for practical workflows. That is a feature of the architecture, not an accident: the core stays small on purpose.
The good news is that the Foclan prompt can be kept stable and cached well, so the prompt-learning overhead is not fully wasted every time. But it is still a real disadvantage today.
Why Give Foclan A Chance
If you are already using Codex, Cursor, Claude Code, or similar tools, Foclan is worth trying for at least three reasons:
- it gives you a concrete way to trade language design against prompt complexity
- it can make exact-output tasks more reliable than plain Python
- it creates a benchmarkable, inspectable middle layer between raw model generation and production code
Even if you never adopt it broadly, Foclan is useful as an experiment in a question that is becoming increasingly practical:
if agents are writing the code, should we keep assuming the optimal language is the same one humans preferred?
What Foclan Is Trying To Become
The goal is not "a universal new programming language".
The goal is a stable and reasonably broad LLM-first language for:
- simple data work
- exact output shaping
- compact intermediate application logic
- practical coding with Codex, Cursor, and similar agent workflows
- data and agent orchestration where each step should stay easy for an LLM to generate correctly
The guiding idea now is:
- Foclan core handles the parts that benefit from being highly constrained
- extensions add practical power without bloating the core
- bridges will eventually provide controlled escape hatches into other runtimes such as Python
In other words: broad enough to be useful, but narrow enough to stay teachable and reliable.
Install
Editable local install:
python -m pip install -e .[test]
Install directly from GitHub:
python -m pip install "git+https://github.com/ovitik/foclan.git"
Optional LLM extension:
python -m pip install "git+https://github.com/ovitik/foclan.git#subdirectory=packages/foclan-llm"
Optional local I/O extension:
python -m pip install "git+https://github.com/ovitik/foclan.git#subdirectory=packages/foclan-io"
Optional HTTP extension:
python -m pip install "git+https://github.com/ovitik/foclan.git#subdirectory=packages/foclan-http"
Optional SQL extension:
python -m pip install "git+https://github.com/ovitik/foclan.git#subdirectory=packages/foclan-sql"
Optional Python bridge runtime:
python -m pip install "git+https://github.com/ovitik/foclan.git#subdirectory=packages/foclan-python"
Build For Publishing
Build the standalone package from the publish/foclan directory:
python -m pip install --upgrade build twine
python -m build
python -m twine check dist/*
This produces the distributable artifacts for the standalone foclan package.
Extension Philosophy
Foclan is intentionally split into:
- a very small core language
- optional extension packages
- and, in the future, bridge runtimes into other languages
Extensions make sense when they do at least one of these:
- simplify common LLM-written workflows
- reduce token count compared with handwritten Python glue code
- reduce failure surface for exact-output tasks
- fit naturally into linear data or agent pipelines
- add practical capability without forcing more syntax into the core language
Extensions do not make sense when they:
- only wrap Python without adding structure or reliability
- introduce many competing idioms
- bloat the prompt teaching burden
- turn Foclan into a second general-purpose language
See also:
Quickstart
foclan examples list
foclan examples validate
foclan examples run counts_dashboard
Start a new local project scaffold:
foclan init project
That writes:
- starter
programs/andinputs/ - a project
README.md .env.exampleAGENTS.md.cursor/rules/foclan-v1.mdc
Render the packaged prompt bundle:
foclan prompt
foclan prompt --anti-overthinking
Optional LLM Extension
The core package stays general and elegant. Provider-specific functionality lives in the optional foclan-llm package.
foclan-llm adds:
.envloading throughfoclan run --dotenvcall llm_textcall llm_json- support for:
- OpenAI Responses API
- Anthropic Messages API
- Google Gemini
generateContentAPI
The extension is intentionally built on the current mainstream provider APIs rather than older legacy endpoints.
Install path:
python -m pip install "git+https://github.com/ovitik/foclan.git"
python -m pip install "git+https://github.com/ovitik/foclan.git#subdirectory=packages/foclan-llm"
Inspect installed extensions:
foclan extensions list
Typical run:
foclan run programs/summarize.focus --env inputs.json --dotenv .env
Bundled LLM examples:
foclan examples list
foclan examples run openai_json_extract --dotenv .env
foclan examples run openai_text_summary --dotenv .env
Notes:
openai_json_extractandopenai_text_summaryrequirefoclan-llm- for text calls, do not starve
max_output_tokens; modern provider APIs may spend part of the budget on reasoning before final text
Optional I/O Extension
foclan-io adds small deterministic file I/O helpers without changing the language core.
It currently provides:
read_textwrite_textread_jsonwrite_jsonread_jsonlread_csvwrite_csv
Install path:
python -m pip install "git+https://github.com/ovitik/foclan.git#subdirectory=packages/foclan-io"
Typical use:
in request
call read_json
out
or:
in request
call write_csv
out
The host function request is just normal data. For example:
{
"request": {
"path": "outputs/report.json",
"content": {"ok": true}
}
}
Optional HTTP Extension
foclan-http adds a very small deterministic HTTP layer as an extension, not as core syntax.
It currently provides:
http_get_jsonhttp_get_texthttp_post_json
Install path:
python -m pip install "git+https://github.com/ovitik/foclan.git#subdirectory=packages/foclan-http"
Typical use:
in request
call http_get_json
out
Headers can be passed directly or backed by environment variables:
{
"request": {
"url": "https://api.example.com/items",
"headers": {
"Authorization": {"env": "API_TOKEN", "prefix": "Bearer "}
}
}
}
This keeps secrets in .env and out of the .focus program itself.
Available Extensions
Current public extensions:
foclan-llmAdds.env-backedllm_textandllm_jsoncalls for OpenAI, Anthropic, and Google.foclan-ioAdds deterministic local text, JSON, JSONL, and CSV file operations.foclan-httpAdds minimal deterministic JSON/text HTTP GET and JSON POST calls.foclan-sqlAdds deterministicsql_queryandsql_execsteps with a SQLite-first request shape.
Current public bridge runtimes:
foclan-pythonAdds a constrained Pythonfocus -> resultbridge for narrow steps that are awkward in pure Foclan.
The intent is not to accumulate random plugins. The intent is to build an ecosystem of extensions that strengthen Foclan specifically for LLM-first data and agent workflows.
Bridging To Other Languages
Bridging is now a public product direction and the first bridge runtime package already exists.
The idea is simple:
- keep Foclan small and optimized for the things it does well
- and provide a controlled escape hatch when a step is better expressed in another language
In practice that means a Foclan program will be able to:
- stay mostly in Foclan for exact shaping and orchestration
- hand the current
focusto another runtime such as Python for one narrow step - then return to Foclan with the new
focus
This changes the role of Foclan in an important way:
- Foclan does not need to become fully general-purpose
- it only needs to be excellent at the LLM-friendly center of the workflow
- and good at handing off the rest in a controlled, low-friction way
The first bridge runtime package is:
foclan-python
The remaining work is to wire the bridge <runtime> ... end syntax fully into the public recommended workflow and extend examples and docs around it.
See:
Public Benchmark
Foclan now ships with a bundled public exact-output benchmark suite.
List the bundled suites:
foclan benchmark list-suites
Run a sampled benchmark against the default Python baseline:
foclan benchmark run \
--provider openai \
--model gpt-5-mini \
--languages foclan python \
--difficulties hard brutal super_brutal \
--sample-size 20 \
--seed 42 \
--reasoning-effort none \
--dotenv .env
The runner writes both JSON and Markdown reports.
Install note:
- benchmark runs need HTTP client +
.envsupport - easiest path is either
foclan-llmorfoclan[benchmark]
Scaffold Editor Integration
In a target project where you want an LLM to write Foclan:
foclan init codex
foclan init cursor
That writes:
AGENTS.mdfor Codex.cursor/rules/foclan-v1.mdcfor Cursor
Repository Layout
- docs/SCOPE.md
- docs/RECOMMENDED_DIALECT.md
- docs/INSTALL.md
- docs/EXTENSIONS.md
- docs/BRIDGES.md
- docs/BRIDGE_SPEC.md
- docs/CODEX.md
- docs/CURSOR.md
- docs/BENCHMARK.md
- docs/FUTURE_CAPABILITIES.md
- prompt
- examples
- templates
- benchmarks
Product Boundary
This repository is intentionally narrow:
- yes: data manipulation, filtering, grouping, sorting, counting, exact response shaping
- yes: compact payload preparation for downstream LLM/API code
- yes: extension-driven capability for file, HTTP, LLM, and future SQL/schema workflows
- yes: future controlled bridging to other runtimes
- no: general-purpose application programming
- no: experimental benchmark harnesses
- no: optimization/search benchmark branches
Roadmap
The near-term plan is:
- stabilize one recommended Foclan dialect
- keep the language reasonably general, but focused on LLM-friendly data work
- reduce the amount of prompt teaching needed
- minimize output tokens and therefore latency
- maximize correctness in both syntax and logic
- improve real usability in Codex and Cursor workflows
- expand capability primarily through extensions
- introduce bridging only if it clearly improves expressiveness without hurting the core simplicity
Only after that foundation is stable should more specialized features be added.
See also:
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file foclan-1.0.0b1.tar.gz.
File metadata
- Download URL: foclan-1.0.0b1.tar.gz
- Upload date:
- Size: 66.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6f264c219ae90cb8195cd2e88f7b1156f53e1bfd46a74c55fb014825d5f5542c
|
|
| MD5 |
84f0a5d7edd0a18af3cd8ba5190ddcda
|
|
| BLAKE2b-256 |
ac4186c09e047a074c7ee06eaccffdb985c3ab04b09f47590c8ebfc5b72f043e
|
File details
Details for the file foclan-1.0.0b1-py3-none-any.whl.
File metadata
- Download URL: foclan-1.0.0b1-py3-none-any.whl
- Upload date:
- Size: 67.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
45c6a722d917633335f5c38680095d64d9b9ec91da1fd18e84773a30d85eb325
|
|
| MD5 |
5846428004ee15b89b35f641d0089363
|
|
| BLAKE2b-256 |
cb8d8ca189bdefa4457b05765a70849958a9077acefc20e38518f671272db0e6
|