Skip to main content

Progress reporting

This page covers two related things: the global graph progress bar (driven by [PROGRESS] / progress events) and the new per-node progress bar for long-running handlers like torch training, monte-carlo loops, and SimWeave's simulate(...).

State of the global progress bar (April 2026)​

The [PROGRESS] token in the worker's stdout protocol is real — the FastAPI bridge in backend/fastapi_main.py parses it and forwards a {type: "progress", value: N} websocket message, and the frontend in workflow.js renders that into the top-level progress bar via updateProgressBar. So the plumbing is wired.

What was not wired is meaningful execution progress. The two places that emitted [PROGRESS] numbers were:

  1. backend/worker_progress.py — a placeholder demo script. Never referenced from production code paths; safe to leave in place as a reference of the protocol shape, or delete.
  2. backend/worker.py end-of-run sweep —
    for i in range(11):
    emit_progress(i * 10)
    time.sleep(0.05)
    This runs after orchestrate_graph_segments returns, so the bar ramps from 0 → 100% in half a second once execution is already done. Pure cosmetics; not coupled to graph progress.
  3. backend/execution/engine.py:427 — inside the loop-segment handler, self.emit_progress(int(100 * (i + 1) / iterations)) does reflect real progress, but only of the loop iteration count, and only for graphs that contain a loop segment.

So Stuart's intuition was correct: for a typical graph (no loop segment) the global progress bar is purely decorative. Any of the following would be a real signal:

  • Nodes-completed / total-nodes: the engine already tracks this implicitly (it emits node_highlight: done per node). Adding a counter and emitting progress = done / total after each completion is a five-line change.
  • Segment-weighted: each segment carries a relative cost (default 1.0); the engine reports progress = sum(weight of done segments) / total_weight. Slightly nicer when one segment dominates the runtime.
  • Wall-clock estimate: keep an exponential moving average of node durations from previous runs, derive an ETA. More work; only worth it for users who run the same graph repeatedly.

Recommendation: start with nodes-completed. It's accurate enough that users will trust the bar without us over-promising. The moving- average path is a future enhancement once we have per-node timing data.

A future PR should replace the cosmetic end-of-run sweep with a single emit_progress(100) to avoid the 0→100 animation that happens after execution has already finished.

Per-node progress bar​

Long-running nodes (torch training, monte-carlo, big SimWeave runs) need their own progress UI inside the node body — the global bar can't disambiguate which node is taking the time, and a graph might have several long-running blocks running concurrently under the executor's worker pool.

Architecture​

handler ── context.report_progress(0.42, "Epoch 5/12") ───┐
▼
engine.emit_node_progress(nid, 0.42, "Epoch 5/12")
▼
worker.emit_node_progress(...)
▼
stdout → fastapi_main.py forwarding layer
▼
websocket {type:"node_progress", node_id, value, label}
▼
workflow.js update_node_progress(...)
▼
<div class="node-progress"> bar inside the node body

Key design choices:

  • Closure-bound node_id. The handler receives a no-arg context.report_progress(value, label=...) — we resolve node_id in the engine's run_single before the call. Handlers stay clean.
  • Fraction (0..1), not percent. Avoids the off-by-one ambiguity that crept into the global bar.
  • Optional label. Free-form short string — handlers can describe what step they're on ("Epoch 5/12", "Particle 4096/10000", "Solver step 1.2s/12s"). The frontend renders it just above the bar.
  • No-op default. context.report_progress defaults to a do-nothing lambda, so existing handlers don't break and new handlers can call it unconditionally.
  • Auto-hide at 100%. The bar fades after a 600ms grace period so the user catches the "done" state without the UI being permanently cluttered.

Handler-side usage​

@register_block("my_long_node", {...})
def my_long_node(node_id, node_def, inputs, context, **kwargs):
n = 100
for i in range(n):
do_work()
context.report_progress(i / n, label=f"Step {i+1}/{n}")
return result, "<pre>done</pre>"

A trivial reference handler lives at projects/demo_project/modules/user_blocks/progress_demo.py. Drop the Loop progress demo node onto the canvas, set Steps and Per-step delay (ms), and run — useful for tuning the bar styling without sitting through a real torch epoch.

Generality vs special cases​

The original question was whether a progress abstraction would generalise beyond "epoch i of N". A few patterns to consider:

Long-running taskNatural fractionNatural label
torch trainingepoch / total_epochs"Epoch {i}/{n}"
torch training (finer)(epoch + batch/n_batches) / total_epochs"Epoch {e}/{E}, batch {b}/{B}"
monte-carlosamples_drawn / total_samples"Sample {i}/{n}"
SimWeave continuous solve(t - t0) / (t1 - t0)"t = {t:.2f}"
SimWeave discrete tick loopcurrent_clock / end_clocksame
HTTP fetchbytes_downloaded / content_length"Downloaded {mb} / {total_mb} MB"
Dataset sweeprows_processed / total_rows"Row {i}/{n}"
Indeterminate (Eg. unbounded API call)n/aspinner

The fraction-and-label abstraction covers everything except the last case. For unbounded work we can introduce a sentinel value (e.g. a negative fraction or None) that the frontend renders as an indeterminate (animated, looping) bar instead of a partial fill — but that's an extension, not the v1.

Torch training integration sketch​

For the torch train_model block, the cleanest hook is a tiny ProgressCallback adapter passed into the train loop:

def train_model(node_id, node_def, inputs, context, **kwargs):
model = ...
opt = ...
epochs = node_def["data"]["epochs"]
for e in range(epochs):
for b, batch in enumerate(loader):
... # forward, backward, step
# Report once per epoch — cheap, and frontend doesn't need
# batch-resolution updates.
context.report_progress(
(e + 1) / epochs,
label=f"Epoch {e+1}/{epochs} loss={running_loss:.4f}",
)
return model, render_summary(history)

If a user later wants batch-resolution updates, swap the inner line for:

context.report_progress((e + b/len(loader)) / epochs, label=...)

Either way the SDK doesn't need a torch-specific abstraction — report_progress is enough.

Cross-cutting concerns​

  • Throttling. A tight inner loop calling report_progress thousands of times per second will flood the websocket and tank the UI. The current frontend transition is 120ms; we should add a debounce on the worker side (e.g. emit at most every 50ms or on > 1% delta) before shipping this to torch's batch loop. Not v1.
  • Concurrency. The engine's _execute_nodes_concurrently runs multiple node handlers in parallel under a thread-pool. Because each closure binds its own node_id and the websocket forwarder is thread-safe, parallel report_progress calls land on the right per-node bar without contention.
  • Stop / interrupt. Handlers that respect stop_requested() can observe it via context if we add the getter — out of scope for the progress doc, captured separately.
  • Tests. Engine unit tests already construct the engine with all emit_* callbacks; adding the optional emit_node_progress keeps them passing because of the no-op default in the engine constructor.

Future work checklist​

  1. Replace the worker's end-of-run 0→100 sweep with a single emit_progress(100).
  2. Wire global progress to nodes-completed / total-nodes in the engine.
  3. Add a debounce/throttle layer on emit_node_progress (50ms or 1%-delta gate) before shipping to torch training.
  4. Indeterminate bar variant for unbounded work.
  5. Per-node ETA from a rolling EMA of historical durations (low priority — only useful for "I run the same graph daily" workflows).