Skip to content

Refactor canvas utility tools for production reliability - #12

Merged
dovvnloading merged 1 commit into
mainfrom
agent/utility-tools-sota-refactor
Jul 15, 2026
Merged

dovvnloading merged 1 commit into
mainfrom
agent/utility-tools-sota-refactor

Conversation

@dovvnloading

Copy link
Copy Markdown
Owner

Summary

  • Fix the QtAwesome frame construction crash and harden Frame/Container membership invariants.
  • Add exact group geometry persistence with stable IDs and preserve collapsed/expanded state on restore.
  • Introduce a shared utility operation lifecycle with bounded context snapshots, cancellation, chat-epoch stale-result guards, provenance, source links, and durable utility saves.
  • Harden Notes with stable metadata, bounded content, undo/redo editing, safe color-picker interactions, and scene-change scheduling.
  • Include utility items in search, highlighting, select-all, and group workflows.

Validation

  • python -m compileall -q graphlink_app
  • python -m pytest -q --disable-warnings --maxfail=1 — 428 passed
  • Headless Qt smoke: real Frame/Container construction, scene rendering, persistence round-trip, and group invariant validation.

The local proposal remains intentionally excluded from the public repository under ignored doc/.

@dovvnloading
dovvnloading merged commit a148e7c into main Jul 15, 2026
6 checks passed
@dovvnloading
dovvnloading deleted the agent/utility-tools-sota-refactor branch July 19, 2026 15:36
dovvnloading added a commit that referenced this pull request Aug 8, 2026
…300)

* ADR-016 stage 16.1: structured JSON logging, run_id, log-level setting

Structured logging: backend/observability.py's JsonLogFormatter renders
one JSON object per line on the rotating file handler (stderr stays
plain text for terminal runs); RunRegistry.claim/cancel/_pop (the one
choke point every one of the app's 12 dispatch surfaces already shares)
now logs "run claimed"/"run cancelled"/"run released" with run_id/kind/
node_id, so every run's lifecycle is traced without touching each
surface individually.

Print-ban: converted the 5 real print() calls in shipped code
(graphlink_settings_store.py x3, graphlink_chat_agent.py,
graphlink_token_estimator.py) to logger calls. Ruff's T201 CI
enforcement is explicitly deferred to ADR-015/roadmap PR #12
("ruff + mypy ratchets") - no ruff config or CI step exists yet in this
repo, and introducing that pipeline is that PR's own scope, not this
ADR's; the underlying state (no print in shipped paths) is met now.

Log-level setting: SettingsManager.get_log_level/set_log_level
(closed vocabulary DEBUG/INFO/WARNING/ERROR, INFO default), a new
setLogLevel app-settings intent that persists AND applies live via
apply_log_level(), and a boot-time read in graphlink_desktop.py's
main() before configure_logging() attaches handlers. Deliberately NOT
wired into backend/app.py's create_app() - that's constructed dozens
of times per pytest run and mutating the root logger's level there
would silently change what caplog captures for unrelated tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* ADR-016 stage 16.2: user-editable pricing + per-node/session cost UI

Backend: SettingsManager.get_pricing_overrides/set_pricing_overrides -
an exact-model-id-keyed local pricing table, checked before the
built-in prefix table in token_counter.py's estimate_cost_usd. Wired
into TokenCounterState via a live accessor (pricing_overrides_fn), so
an override change applies to the next reply without a restart.

ChatState gains estimated_cost_usd - a point-in-time cost snapshot
taken when a reply's usage is stamped (backend/api/intents_chat.py's
_on_usage), surviving save/load like prompt_tokens/completion_tokens.
TokenCounterState gains session-cumulative totals (sessionPromptTokens/
sessionCompletionTokens/sessionEstimatedCostUsd), accumulated on every
real-usage reply.

Frontend: ChatNodeView renders a compact "prompt→completion · $cost"
badge next to the model badge - data was on the wire since ADR-006
stage 6.8 but never rendered. No settings UI for editing pricing
overrides in this sweep, matching the get_mcp_servers/set_mcp_servers
precedent from ADR-007 stage 7.5 (data layer now, UI panel deferred).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* ADR-016 stage 16.3: in-app diagnostics (runs, publish size/rate, provider errors)

backend/diagnostics.py adds a per-session DiagnosticsState (recent run
outcomes/durations, publish count/bytes/rate, session count) fed by explicit
on_claim/on_end callbacks on RunRegistry and a publish-size recorder on
SessionBus, plus a process-global bounded deque of provider errors recorded
from api_provider.py's single chat-exception translation choke point. Wired
into app.py as a new "diagnostics" WS topic.

Frontend: DiagnosticsDialog.tsx subscribes to the topic and renders the
stats, a recent-runs table, and a provider-errors list, reachable from the
app bar next to About/Help. Also fixes a pre-existing CSS overflow bug in
the token-counter popout (unbounded width from ADR-006 stage 6.8 growing it
to 7 columns without revisiting the layout).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* ADR-016 stages 16.4 + 16.5: diagnostic bundle, open-log-folder, evals harness

Stage 16.4 - backend/diagnostic_bundle.py builds a redacted snapshot
(app version, OS, node counts by kind, the diagnostics payload) for issue
reports. Redaction is structural: build_diagnostic_bundle's signature takes
only (document, diagnostics_state), so no settings manager, log handle, or
provider object is reachable from inside it. Exposed as two read-only
intents on the existing "diagnostics" topic, with an action row in the
Diagnostics dialog that previews and copies the bundle.

One field was not content-free by construction: providerErrors[].message
started as raw str(exc), which embeds the full absolute path (including the
OS username) for an OSError and can carry SDK-echoed request content.
record_provider_error now classifies it into a fixed vocabulary of category
labels and stores no substring of the original; the full text still reaches
graphlink.log.

Stage 16.5 - backend/evals/ runs golden fixtures against the chart and
structured-output paths, deterministic by default (provider seam patched)
and --live against a real provider. The agent-build dimension reports
not_implemented rather than a fake pass: ADR-008's Builder loop does not
exist yet, so that half of the stage's exit criterion is not met and the
ADR now says so.

Test plan: full pytest from repo root (2045 passed, 14 skipped); web_ui
typecheck, lint, test (1382 passed), build, and codegen drift check all
clean; python -m backend.evals exits 0 with 4 passed / 1 not_implemented.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant