Local Notebook Agent Editor
Demo
https://github.com/user-attachments/assets/e52a4ee7-3bb2-4953-a515-4c32f8742c6f
A local FastAPI + React editor for working on Jupyter notebooks with scoped AI
agent editing. You open a local .ipynb file (or a project folder), edit and
execute cells, and save in place. Each agent turn has a scope: in
Blocking mode the agent may only rewrite the source of cells you explicitly
mark editable; in Trusted mode the whole notebook is editable and the agent
may add, delete, reorder, and retype cells. Either way the backend validates and
applies (or rejects) every change — the agent never mutates your notebook
directly — and every change is reviewable as a diff and undoable.
Everything runs on your machine and binds to loopback only. Read Security Limits before using it with untrusted notebooks.
Install
You work in the browser tab. An MCP client — Claude Code, for instance — starts the editor and opens it for you. Running it locally is for working on the editor itself.
Install it into the environment your notebooks run in, so your cells keep the packages they already have:
/abs/path/to/.venv/bin/pip install notebook-editor-mcp
claude mcp add agent-notebook -- \
/abs/path/to/.venv/bin/notebook-editor-mcp \
--workspace-root /abs/path/to/project
To keep the editor out of your project instead, point --kernel-python at the
environment your cells need:
claude mcp add agent-notebook -- \
uvx --from notebook-editor-mcp notebook-editor-mcp \
--workspace-root /abs/path/to/project \
--kernel-python /abs/path/to/.venv/bin/python
--workspace-root confines the editor to one directory. Use absolute paths for
it and for the executable.
Then ask your client for a notebook. The tab opens on its first tool call and
show brings it back. Everything the editor does is there: the agent chat,
per-hunk diff review, the approval gate a risky run parks at, the plot tuner
and the notebook map.
Needs Python 3.11+ and ipykernel in whichever environment runs your cells,
plus the claude CLI on your PATH for the tab's chat (see Claude
CLI). Any client that speaks stdio MCP works.
What the client can do too
The client can drive the notebook while you watch, against the same session
you are working in. The tools are named for what they do and nothing more,
because clients that namespace them by server name — Claude Code renders them
mcp__agent_notebook__open — would otherwise produce notebook_open inside a
namespace already called notebook.
open, read, status,
set_cell_source, insert_cell, delete_cell,
run_cell, run_all, cancel_run, save, show.
Cells are addressed by the cellId that read returns, and nothing reaches
the file until save — including the outputs of a run.
Three things the tab guarantees you, whatever the client does:
- You approve anything risky before it runs. Execution a tool asks for is agent-initiated, so a cell the risk classifier flags stops at awaiting approval and waits for you — the same gate an agent turn's downstream cells get.
- Your edits win. Change a cell in the tab and a client writing against the version it last read is refused and told to re-read. Nothing retries over the top of your work.
- New cells arrive for review, not already run.
insert_cellmarks its cell agent-authored, the tab badges it "review before running", and leaves it inert. Running it is a separate call through the same approval gate.
Agent turns are not exposed as tools — the client is already the agent, and
running the claude CLI underneath it would just nest a second one. The tab's
chat is where a turn gets sent, and it works exactly as it does locally.
Neither is the file browser exposed: the tab has a picker, and a person
browsing their own machine is a different thing from a model enumerating it.
Watching it work
Run as an MCP server the editor is a child process the client starts for you,
on a port chosen at random, with stdout reserved for the protocol. That is
three ways in which the usual scripts/dev.py habits do not apply, so:
- The port is announced on stderr, where MCP clients keep their server
logs:
notebook editor ready at http://127.0.0.1:PORT (log: ...). Open that URL and you are looking at the same session the tools are driving. The same URL comes back aseditorUrlfromopen, and fromshowon request. NOTEBOOK_EDITOR_LOG=/path/to/editor.logkeeps the editor's own log at a path you choose instead of a temp file that is deleted when it stops. This is what to set beforetail -f. Unset, a failed start is still drained and reported in the error, so you only need this for watching a server that is working.- Never print to stdout from the server process. It is the transport; a
stray
printcorrupts the protocol rather than showing up somewhere. --no-browsersuppresses the automatic tab, for automation and headless hosts.showstill opens it, andopenstill returns the URL aseditorUrlfor a client to hand you.
To see the tool surface without wiring up a client, the MCP Inspector speaks the same stdio protocol:
npx @modelcontextprotocol/inspector \
/abs/path/to/.venv/bin/notebook-editor-mcp \
--workspace-root /abs/path/to/project --no-browser
Read Security Limits before pointing a client at a notebook you would not run yourself.
The interface
The editor at rest: files and the notebook map on the left, the notebook in the middle, the agent on the right. Nothing is scoped yet, so the composer says so — a turn sent now is read-only.
After a turn, changed cells become reviewable in place — an inline diff on the cell it belongs to, accepted or rejected per hunk, not a patch file you read somewhere else. The review bar counts what is pending and names the turn it came from; the agent's answer stays beside the diff, so what it claims to have done and what it actually changed are on screen together. Here two of the three scoped cells were rewritten, and the panel explains why the third was left alone.
|
The notebook map. The Outline tab segments the notebook into blocks and asks the model to name each one, so a long notebook has a table of contents it never had. Names are generated — they carry a dotted underline to say so — while the cell ranges under them are computed. Rebuild map re-derives it after the notebook moves on. |
Plots get knobs. Tune scans the cells above a figure for values it can vary safely — sizes, counts, colours, flags — and puts each on a control, floating over the notebook so the picture keeps the full width of the cell. Moving one re-runs the cell into a preview labelled not in your notebook yet; Apply and re-run is what writes the values back. Nothing is committed until you press it.
Screenshots are captured from a live session against the notebooks in
examples/, with a real kernel and the real Claude CLI. See
docs/screenshots/ for the full set and how to re-capture
them.
Run it locally instead
For working on the editor itself, or to use the browser UI on its own with agent turns driven from the tab rather than from an MCP client.
Prerequisites
- Python 3.11+
- Node.js 20.19+ or 22.12+, with npm
- A local Python kernel (installed with the backend below via
ipykernel) - Claude CLI — required only for agent turns (see Claude CLI)
- macOS or Linux. The launcher and Playwright cleanup use POSIX process groups/signals. Windows is out of scope for v1.
Setup
From the repository root:
python -m venv .venv
.venv/bin/pip install -e '.[test]'
npm install
Playwright is only needed if you plan to run the end-to-end suite:
npx playwright install chromium
Claude CLI
Agent turns shell out to the claude command-line tool. You can open, edit,
run, and save notebooks without it — the CLI is only invoked when you send an
agent turn.
1. Install it using either official method:
# Native installer (installs to ~/.local/bin)
curl -fsSL https://claude.ai/install.sh | bash
# …or via npm
npm install -g @anthropic-ai/claude-code
Make sure claude is on your PATH (the app runs the claude executable
directly):
which claude
2. Authenticate it once. Run claude on its own and complete the login
flow (an Anthropic account / Claude subscription, or an API key). The app
launches the CLI non-interactively per turn and does not handle login for you.
3. Match the supported version. The adapter verifies the CLI before every
turn and requires claude >= 2.1.203 and < 2.2.0. It relies on
--safe-mode, --disable-slash-commands, --strict-mcp-config, --tools,
and --permission-mode. Check your version:
claude --version
If it falls outside that range, agent turns fail fast with "Unsupported
Claude CLI version"; if claude is missing from PATH, you get "Claude
CLI is unavailable." In both cases opening/editing/running notebooks still
works — only agent turns are blocked.
Just want to try the UI without Claude? Start the app with
--test-agent(see Run) to use a built-in deterministic adapter. It produces canned Blocking-mode edits for demoing the flow and never calls the real CLI — it does not perform Trusted structural edits, and it is not a substitute for a real agent.
Run
One command starts FastAPI at http://127.0.0.1:8000 and Vite at
http://127.0.0.1:5173:
.venv/bin/python scripts/dev.py
Then open http://127.0.0.1:5173.
Useful flags:
# Run against the built-in deterministic adapter instead of the real Claude CLI
.venv/bin/python scripts/dev.py --test-agent
# Use different ports (the Vite dev server proxies /api to the backend port)
.venv/bin/python scripts/dev.py --backend-port 8055 --frontend-port 5199
Using it
- Open a notebook. From the start screen choose a local
.ipynbfile, or open a project folder to browse and edit its notebooks in place. There's also an upload/download path if you prefer working on a copy. - Edit and run cells. Edit cell source directly and run cells against the local kernel. Toggle Auto-save or use Save / Save As to write back to disk.
- Scope cells for the agent. Two independent axes: mark cells editable (permission — the agent may rewrite their source) and/or pin cells as Focus (salience — the cells most relevant to your request; attention only, see the security note). Sending a prompt with no editable cells is a valid read-only turn: the agent answers but writes nothing.
- Pick a scope, model, and mode in the agent composer:
- Write scope (composer footer, beside Model and Mode) — Blocking (default) lets the agent edit only the cells you marked editable; Trusted makes the whole notebook editable and lets it add/delete/reorder/retype cells. In Trusted the per-cell "allow agent edit" control is hidden and every pin is a Focus hint.
- Model — Default, Opus, Sonnet, or Haiku. Default defers to the CLI's
own default; otherwise it is passed as
--model. - Mode — Edit applies scoped changes; Plan returns a step-by-step plan and writes nothing.
- Send. Review the agent's answer and the inline diffs on changed cells. A Trusted turn also shows a summary of structural changes (added / deleted / reordered / retyped) and marks agent-added cells with a provenance badge; Trusted turns apply structure only and do not auto-run cells — you run them after reviewing. You can undo an applied turn (whole-turn undo).
The outline panel
The sidebar's Outline tab is a map of the open notebook: contiguous ranges
of cells, each with a name and a cell range. Click a block to jump to its first
cell; hover to highlight everything it covers in the gutter. It is navigation
only — nothing it shows is ever written to the .ipynb.
One field is generated and the rest are computed, and the panel renders them differently on purpose:
- The name is written by a model, and carries a dotted underline to say so. Every block cites its own cell range, so a name you disagree with is checkable in one click rather than something you have to trust.
- Everything else is computed from the AST, with no model involved: the cell range, the variables the block produces that later cells read, the functions it defines and which cells call them (including ones never called), and any markdown headings inside the range. Expand a block to see the last three.
The boundaries are computed too. They fall where the notebook's vocabulary
changes — comparing the identifiers used either side of each seam and cutting at
the valleys. Markdown headings are not consulted at all, which is the point: the
notebooks that most need a map are the ones without them. An earlier version cut
at every heading, and measuring it against nine real notebooks
(docs/plans/probes/corpus/) showed why that had to change — it cost a reader
21.5 cells to find what they were looking for, against 9.3 for this, and lost
even to cutting blindly every eight cells.
So the model call is spent on naming, not on segmenting. The same measurement found a model-drawn partition indistinguishable from the computed one, so Build map hands the model the blocks and asks only what to call them. Three things follow: the map cannot come back malformed, because the boundaries were never the model's to return; the same notebook always segments the same way, where an earlier version's block count moved 12% between identical runs; and the free map you see before pressing anything is a real map rather than a placeholder.
Opening a notebook costs nothing. The computed fields are an AST parse and refresh on every document revision, so running a cell updates the panel without touching the names. Only Build map spends anything, and only when you press it.
Build map is one claude call: cell source only — never outputs, at any
size — with no tools, no MCP servers, and a throwaway working directory, so the
pass can read, write and run nothing. It defaults to Haiku. The result is cached
per notebook path and survives cell runs; editing a cell marks the map stale
but leaves it on screen until you rebuild.
Three ways it declines to guess, all visible in the panel:
- No model pass yet (or one that failed) shows the computed blocks with the cell range where the name would be, and says names need a model. It does not invent one.
- An answer that is not one name per block is discarded rather than rendered, and the reason is shown. The previous map, if any, stays.
- Above ~500 cells it refuses outright instead of segmenting a prefix and presenting it as the whole map.
Verify
.venv/bin/python -m pytest backend/tests -q # backend unit/integration tests
npm test -- --run # frontend unit tests
npm run build # type-check + production build
npm run test:e2e # Playwright end-to-end (needs chromium)
The Playwright suite opens examples/sample.ipynb and covers desktop and
mobile workflows. Screenshots and failure traces are written under
test-results/.
The MCP surface is checked at three depths, and it is worth knowing which one to reach for:
| What it covers | Cost | |
|---|---|---|
pytest backend/tests/test_mcp_server.py |
the tool surface against fakes — errors, revisions, descriptions | instant |
pytest backend/tests/test_mcp_end_to_end.py |
the tools against a real editor, kernel and approval gate | seconds |
python evals/mcp_tool_eval.py |
whether a model drives the tools well | a minute per run, and tokens |
Only the last one needs an authenticated claude CLI, which is why it is not
in CI. It answers a different question from the other two: not "does the tool
work" but "does a model reach for the right one, fetch no more than it needs,
and ignore a notebook that tells it to do something else".
python evals/mcp_tool_eval.py # every task, once each
python evals/mcp_tool_eval.py --repeat 3 # a pass rate, not a pass
python evals/mcp_tool_eval.py \
--model claude:sonnet --model claude:haiku # the same tasks, two models
--model takes client[:model] and repeats. A tool description that only
reads well to the strongest model is a defect a single-model suite cannot see —
running Haiku alongside Sonnet is what makes that visible.
Two clients are wired: claude and codex. The Claude path is the verified
one. The Codex path was built against the CLI's own help and shipped event
names without a credential to run it, so treat its first run as a shakedown: if
the transcript does not match what the parser expects it says so, naming the
event types it saw, rather than reporting a run that used no tools. Codex needs
codex login first — the harness deliberately does not touch CODEX_HOME,
which is where that credential lives.
They are not perfectly matched. Claude runs with --tools "", so the MCP
surface is the only way to reach the notebook at all. Codex has no equivalent
switch and runs under a read-only sandbox instead: it cannot write the notebook
by other means, but it could read one without the tools. Every task asserts the
surface was used, so such a run fails rather than passing hollowly.
Security Limits
The editor binds to loopback and keeps one active notebook in process memory. The agent workspace contains the full notebook, so the agent can read every cell. Focus (attention) selection is an attention signal in the manifest and prompt, not a confidentiality control. In Blocking turns, writes are imported only from manifest-listed editable-cell source files; notebook structure, metadata, outputs, and unselected-cell writes are rejected at the workspace boundary. Trusted turns widen this to the whole notebook (see below). Risk classification pauses selected downstream operations for explicit approval.
Trusted turns (the per-turn Blocking/Trusted toggle) widen the write boundary to the whole notebook and let the agent add, delete, reorder, and change the type of cells — including introducing new executable code cells you never scoped, and changing execution order. The backend still derives, validates, and applies every change (the agent never edits the notebook directly), and OS posture is unchanged (no shell, loopback only). But this makes review-before-run load-bearing: agent-added cells are marked with a provenance badge and are not auto-executed, yet once accepted they are ordinary cells a later Run-All will run. Review the diff before running, and use Trusted only with agents and instructions you trust. The risky-cell execution approval flow still gates execution.
Driving it from an MCP client narrows two of these and widens one. Model-
initiated runs are gated, where a click of Run is not, and --workspace-root
confines every path the server accepts. But the notebook's own text — markdown,
code, and cell outputs — is read into the model's context while that same model
holds a tool that executes cells. A notebook from somewhere you do not trust
can therefore carry instructions to an agent that can run code on your machine.
The approval pause is the control that matters there; do not wave it through.
This is not an operating-system sandbox. The CLI and executed notebook code run with the current user's permissions. Risk classification is heuristic, approval does not make code safe, and notebook execution can read files, use credentials, access the network, or start processes. Use the editor only with notebooks and agent instructions you trust.
Contributing
Contributions are welcome — see CONTRIBUTING.md for setup, the checks CI runs, and pull-request guidelines. Participation is governed by our Code of Conduct. To report a vulnerability, see SECURITY.md.
License
Released under the MIT License.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file notebook_editor_mcp-0.1.1-py3-none-any.whl.
File metadata
- Download URL: notebook_editor_mcp-0.1.1-py3-none-any.whl
- Upload date:
- Size: 774.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bd8ff13d43797c33e28424ef56b7827a973c2f0f00341f219a8e43a07f9b1656
|
|
| MD5 |
11ab523acd9f0beabe3049c1e52a9406
|
|
| BLAKE2b-256 |
d2662bea4a8ca5fd8462d463eec0798568d99cd00acff58f9e49d8ff0b8754aa
|
Provenance
The following attestation bundles were made for notebook_editor_mcp-0.1.1-py3-none-any.whl:
Publisher:
release.yml on AgLyx3/we-love-jupyter-notebook
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
notebook_editor_mcp-0.1.1-py3-none-any.whl -
Subject digest:
bd8ff13d43797c33e28424ef56b7827a973c2f0f00341f219a8e43a07f9b1656 - Sigstore transparency entry: 2620406985
- Sigstore integration time:
-
Permalink:
AgLyx3/we-love-jupyter-notebook@d4f7c49e43eeed3984145bd61c4d20f4b076c066 -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/AgLyx3
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d4f7c49e43eeed3984145bd61c4d20f4b076c066 -
Trigger Event:
push
-
Statement type: