Skip to main content

Jupyter notebook viewer — design note (deferred)

Intent​

EdgeWeave is a Python-first workflow tool, and a non-trivial number of its likely users will already have notebooks lying around — exploratory analyses, scratchpads, course material, a colleague's prototype. Letting users open a .ipynb in the EdgeWeave code editor (with no ability to execute it) is a high-leverage, low-risk feature: it makes EdgeWeave a more natural place to drag old work into, without dragging in the security and dependency surface of a kernel.

This doc captures the shape so the eventual PR can move directly to implementation. No code yet.

Hard constraint: read-only, no execution​

  • No kernel, no ZeroMQ, no nbformat execution.
  • No write-back to the .ipynb from the editor — opening a notebook produces a viewer pane, not an editable buffer. Users wanting to edit cells should "Convert to .py", which extracts code cells into a fresh Python file.
  • No "trusted notebook" prompt. Outputs render through a sandboxed HTML pipeline that allows known-safe MIME types only.

This is the smallest feature set that's actually useful, and it sidesteps every notebook security pitfall (untrusted JS in display_data, kernel escape via maliciously-crafted ipynb metadata, etc.).

File structure​

.ipynb is JSON. The shape we care about is small:

{
"cells": [
{
"cell_type": "code" | "markdown" | "raw",
"source": ["line 1\n", "line 2\n"] // or a single string
"outputs": [ ... ] // code cells only
}
],
"metadata": { "kernelspec": { ... } } // ignored
}

Output entries (code cells only) come in four types:

  • stream — {name: "stdout"|"stderr", text: "..."}
  • execute_result / display_data — {data: {<mime>: <value>}, ...}
  • error — {ename, evalue, traceback: ["...ANSI..."]}

We render the first three; error cells render with a red-tinted background and the traceback rendered through a small ANSI-to-HTML helper (existing terminal pane already has one).

Phasing​

v1 — core viewer (~200–300 LOC, ~1 focused day)​

  • Accept .ipynb in the existing file-open path.
  • Backend route GET /api/notebook?path=... returns the parsed JSON (or 400 on malformed).
  • New frontend/src/js/notebook_viewer.js:
    • Renders cells in document order into a scrolling pane.
    • Code cells → read-only CodeMirror 6 view, language defaulted to Python (python() is already imported). Reuses the same CodeMirror infrastructure the chat code-block rendering uses (code_languages.js).
    • Markdown cells → run through the same marked + DOMPurify pipeline the chat assistant uses. Free reuse, free LaTeX support if we ever add KaTeX for the chat too.
    • Raw cells → <pre> block.
  • Output rendering — render only these MIME types in v1:
    • text/plain → <pre>
    • image/png / image/jpeg → <img src="data:..."> with a max width / height to prevent gigantic notebook images from breaking layout.
    • text/html → through DOMPurify with a strict allow-list (no <script>, <iframe>, <object>, <embed>).
    • application/json → pretty-printed <pre>.
    • Any other MIME type → a "rendered output omitted (type: X)" placeholder. Better than guessing wrong.
  • Default styling matches the current dark/light theme via the same CSS variables (--panel-bg, --text-color, etc.).

v2 — polish (later, optional)​

  • Plotly/Altair JSON output rendering — both ship browser bundles that can render a application/vnd.plotly.v1+json payload directly. Bundle cost is ~3MB, so this should be lazy-loaded only when a notebook actually contains a plotly output.
  • Cell folding (collapse long output blocks).
  • "Convert to .py" — drop a fresh <file>.py next to the notebook containing only the code cells, joined with # %% markers so the Jupytext "percent" round-trip works.
  • Search-within-notebook.

Why not "just display in CodeMirror as JSON"​

Tempting, since .ipynb is JSON and we'd get language-multipack dispatch for free. But:

  1. The actual content (code, markdown, plotted images) is what the user wants to see, not the cell metadata. A raw JSON view buries the signal.
  2. Outputs serialised in JSON look like "data": {"image/png": "iVBORw0KGgo..."} — useless to the eye.

So we want a small dedicated viewer.

Plumbing reuse​

The v1 viewer leans on existing pieces, which is why it's only ~1 day of work despite sounding like a feature:

ReuseSource
Markdown → safe HTMLfrontend/src/js/chat_ui.js (M2.5 patch)
Read-only CodeMirror with themefrontend/src/js/codemirror_main.js
Language dispatch (.py → python() etc.)frontend/src/js/code_languages.js
File picker / open-in-editor IPCexisting
ANSI → HTML for error tracebacksterminal pane

Risks​

  • Notebook size — a notebook with hundreds of large image outputs can easily be >50 MB. The HTTP fetch is fine; the DOM tree built from rendering all of it at once may not be. Plan: viewport-based lazy rendering (render cells as they scroll into view) if the file exceeds a size threshold (~5 MB). For v1 we can ship without this and accept that 50 MB notebooks render slowly.
  • Wide tables in text/html outputs — pandas often emits huge HTML tables. Allow horizontal scroll on .notebook-output table via overflow-x: auto and a max-width: 100% wrap.
  • DOMPurify allow-list drift — the strict allow-list will reject some cells that would have rendered fine (e.g. notebook authors who include <details> or <summary>). Add to the allow-list as legitimate cases come up; do not auto-relax.

When to do this​

After:

  1. M2.5 markdown rendering ships (this provides the markdown pipeline the viewer reuses).
  2. CodeMirror multi-language dispatch ships (provides the language resolution the viewer reuses).

Both of those are landing in the chat-backend follow-up branch (feature/chat-backend), so the notebook viewer can come right after on its own branch.

Out of scope (do not let this creep)​

  • Notebook editing. If users want that, they want JupyterLab.
  • Kernel execution. Same.
  • Notebook creation. Use the existing .weave workflow; export to .ipynb at the end if needed.
  • Reading remote notebooks (Colab, GitHub). Local files only for v1.