Jupyter notebook viewer — design note (deferred)
Intent
EdgeWeave is a Python-first workflow tool, and a non-trivial number of
its likely users will already have notebooks lying around — exploratory
analyses, scratchpads, course material, a colleague's prototype.
Letting users open a .ipynb in the EdgeWeave code editor (with no
ability to execute it) is a high-leverage, low-risk feature: it makes
EdgeWeave a more natural place to drag old work into, without dragging
in the security and dependency surface of a kernel.
This doc captures the shape so the eventual PR can move directly to implementation. No code yet.
Hard constraint: read-only, no execution
- No kernel, no ZeroMQ, no
nbformatexecution. - No write-back to the
.ipynbfrom the editor — opening a notebook produces a viewer pane, not an editable buffer. Users wanting to edit cells should "Convert to .py", which extracts code cells into a fresh Python file. - No "trusted notebook" prompt. Outputs render through a sandboxed HTML pipeline that allows known-safe MIME types only.
This is the smallest feature set that's actually useful, and it sidesteps every notebook security pitfall (untrusted JS in display_data, kernel escape via maliciously-crafted ipynb metadata, etc.).
File structure
.ipynb is JSON. The shape we care about is small:
{
"cells": [
{
"cell_type": "code" | "markdown" | "raw",
"source": ["line 1\n", "line 2\n"] // or a single string
"outputs": [ ... ] // code cells only
}
],
"metadata": { "kernelspec": { ... } } // ignored
}
Output entries (code cells only) come in four types:
stream—{name: "stdout"|"stderr", text: "..."}execute_result/display_data—{data: {<mime>: <value>}, ...}error—{ename, evalue, traceback: ["...ANSI..."]}
We render the first three; error cells render with a red-tinted
background and the traceback rendered through a small ANSI-to-HTML
helper (existing terminal pane already has one).
Phasing
v1 — core viewer (~200–300 LOC, ~1 focused day)
- Accept
.ipynbin the existing file-open path. - Backend route
GET /api/notebook?path=...returns the parsed JSON (or 400 on malformed). - New
frontend/src/js/notebook_viewer.js:- Renders cells in document order into a scrolling pane.
- Code cells → read-only CodeMirror 6 view, language defaulted
to Python (
python()is already imported). Reuses the same CodeMirror infrastructure the chat code-block rendering uses (code_languages.js). - Markdown cells → run through the same
marked+DOMPurifypipeline the chat assistant uses. Free reuse, free LaTeX support if we ever add KaTeX for the chat too. - Raw cells →
<pre>block.
- Output rendering — render only these MIME types in v1:
text/plain→<pre>image/png/image/jpeg→<img src="data:...">with a max width / height to prevent gigantic notebook images from breaking layout.text/html→ throughDOMPurifywith a strict allow-list (no<script>,<iframe>,<object>,<embed>).application/json→ pretty-printed<pre>.- Any other MIME type → a "rendered output omitted (type: X)" placeholder. Better than guessing wrong.
- Default styling matches the current dark/light theme via the same
CSS variables (
--panel-bg,--text-color, etc.).
v2 — polish (later, optional)
- Plotly/Altair JSON output rendering — both ship browser bundles
that can render a
application/vnd.plotly.v1+jsonpayload directly. Bundle cost is ~3MB, so this should be lazy-loaded only when a notebook actually contains a plotly output. - Cell folding (collapse long output blocks).
- "Convert to .py" — drop a fresh
<file>.pynext to the notebook containing only the code cells, joined with# %%markers so the Jupytext "percent" round-trip works. - Search-within-notebook.
Why not "just display in CodeMirror as JSON"
Tempting, since .ipynb is JSON and we'd get language-multipack
dispatch for free. But:
- The actual content (code, markdown, plotted images) is what the user wants to see, not the cell metadata. A raw JSON view buries the signal.
- Outputs serialised in JSON look like
"data": {"image/png": "iVBORw0KGgo..."}— useless to the eye.
So we want a small dedicated viewer.
Plumbing reuse
The v1 viewer leans on existing pieces, which is why it's only ~1 day of work despite sounding like a feature:
| Reuse | Source |
|---|---|
| Markdown → safe HTML | frontend/src/js/chat_ui.js (M2.5 patch) |
| Read-only CodeMirror with theme | frontend/src/js/codemirror_main.js |
Language dispatch (.py → python() etc.) | frontend/src/js/code_languages.js |
| File picker / open-in-editor IPC | existing |
| ANSI → HTML for error tracebacks | terminal pane |
Risks
- Notebook size — a notebook with hundreds of large image outputs can easily be >50 MB. The HTTP fetch is fine; the DOM tree built from rendering all of it at once may not be. Plan: viewport-based lazy rendering (render cells as they scroll into view) if the file exceeds a size threshold (~5 MB). For v1 we can ship without this and accept that 50 MB notebooks render slowly.
- Wide tables in
text/htmloutputs — pandas often emits huge HTML tables. Allow horizontal scroll on.notebook-output tableviaoverflow-x: autoand amax-width: 100%wrap. - DOMPurify allow-list drift — the strict allow-list will reject
some cells that would have rendered fine (e.g. notebook authors
who include
<details>or<summary>). Add to the allow-list as legitimate cases come up; do not auto-relax.
When to do this
After:
- M2.5 markdown rendering ships (this provides the markdown pipeline the viewer reuses).
- CodeMirror multi-language dispatch ships (provides the language resolution the viewer reuses).
Both of those are landing in the chat-backend follow-up branch
(feature/chat-backend), so the notebook viewer can come right after
on its own branch.
Out of scope (do not let this creep)
- Notebook editing. If users want that, they want JupyterLab.
- Kernel execution. Same.
- Notebook creation. Use the existing
.weaveworkflow; export to.ipynbat the end if needed. - Reading remote notebooks (Colab, GitHub). Local files only for v1.