Skip to content

Latest commit

 

History

History
770 lines (579 loc) · 51.2 KB

File metadata and controls

770 lines (579 loc) · 51.2 KB

Changelog

All notable changes to Cortex are recorded here, newest first, in the Keep a Changelog layout. The number of the pull request behind a change is given in parentheses.

The declared version is 2.0.0 (backend/cortex_backend/__init__.py). It is not published: there is no download, so run Cortex from source. The newest published release is v1.0.0, from the Qt-era application that the web rewrite replaced. Nothing after v1.0.0 (2026-01-20) has been published, so both the summary below and the dated entries under it are unreleased. When a release is cut, [Unreleased] becomes ## [<version>] - <date>.

[Unreleased]

Everything merged since 2026-09-04 (pull requests #212 to #291), which this file had no summary of. Test-only pull requests (#228, #269, #273) and a README edit that was later undone (#278) are left out. Some changes are also told at length in the dated entries under Earlier entries, which are named where they apply.

Added

  • Answers stream into the transcript as the model writes them, reasoning included, instead of appearing all at once after generation (#268; see "Real Token Streaming" below).
  • A redesigned model picker and generation-parameter controls (#271).
  • Models kept in subfolders of the GGUF folder are found, and a model's id carries its path (#270; see "GGUF Model Discovery" below).

Changed

  • The declared version is 2.0.0, marking the rewrite as a different application from the Qt-era v1.0.0. It is not published (#277).
  • The release workflow builds the tag it was asked to build, and only its upload step can write to the repository (#277).
  • Startup is faster: Cortex builds one TLS context and shares it between its three HTTP clients instead of parsing the certificate bundle three times (#284).
  • The README describes the app as it ships: both runtimes, screenshots of the current build, and what code execution can reach today (#276).
  • Internal restructuring with no change in behaviour: the chat store is named for what it is (#212), the composition root is app_factory (#214), each local worker attempt has its own module (#221), execution capability gates are typed (#223), the services call the generation engine instead of interrogating it (#226), and the llama.cpp chat stream is stopped by closing it rather than through a queue (#227).
  • For contributors: a Python type checker (mypy) is part of the checks and the package's py.typed marker now holds (#225); tests are named for what they cover (#220), a browser spec runs against the real API (#213, #216), tests that never touch a DOM no longer build one (#229), every wait in the suite is bounded (#282), and the GitHub Actions are pinned to commits, the workflows are linted and audited, and pull requests get a dependency review (#283). Tests share fixtures, a frozen repository clock and a check that no worker thread outlives the run (#285), and the repository documents were brought back in step with the repository (#288).

Fixed

Chat and streaming

  • Stop keeps the part of the answer already on screen instead of erasing it (#279; see "Stop Keeps the Answer" below).
  • Stop is honoured during translation (#252), and a turn that was already cancelled no longer waits for the model (#242).
  • A failed translation no longer breaks the live answer stream (#274; see "Translation Failures" below).
  • Nine defects in cancellation, consent, time and startup (#266; see "Correctness Pass" below).
  • A prompt such as "start the process of writing" is no longer treated as a request for code (#281).
  • The transcript settles when the reload after a generation fails (#235), and a turn's status callback is detached when the turn ends (#264).
  • An image is refused for a GGUF model that cannot see it, instead of being dropped (#234).
  • A document in an unknown text encoding is refused rather than guessed at (#250), and the image capability probe no longer breaks GIF attachments (#248).
  • An ordinary chat message can no longer freeze the backend (#247).
  • Replies are parsed more carefully: an unfinished memory or code block no longer leaks into the answer, an identical memory block counts once, quoted examples are left alone, a leading <think> block is shown as reasoning, and command tags hidden in attachments, memories or tool output are neutralised (#291).
  • An offline runtime and a timed-out model get their own messages instead of the generic failure, and the system message keeps the same beginning from turn to turn so the model's cache is reused (#291).

Interface

  • Retry is offered only for failures that resending the prompt can fix (#263), and the composer clears when a retry succeeds (#265).
  • A local generation keeps streaming while the machine is offline (#258).
  • A cleared number field is no longer read as zero (#262).
  • Unsaved memory edits survive when a memory is added (#261).
  • Attachments that uploaded are kept when another one in the same batch fails (#260).
  • Three state bugs that threw away visible work: a stale 401 no longer signs the user out of a working session, a failed chat move no longer undoes a later one that succeeded, and memories the server normalized away no longer stay on screen (#241).
  • An expired session is renewed in place instead of reloading the workspace, a generation keeps running when its stream meets an expired session, and a failing memory store no longer blocks the whole workspace (#290).

Code execution

  • A cancelled job no longer reports success (#239), cancelling twice stays idempotent (#267), and a cancelled code task can no longer run to completion (#266).
  • Two lifecycle races that cost work: a duplicate submission no longer kills the job it duplicated, and reopening Cortex soon after a crash no longer leaves code execution disabled for the whole session (#233).
  • A damaged execution store is rebuilt instead of refusing to start (#240).
  • The execution event stream is never closed before its terminal event (#231), and events are no longer evicted from under a reader that is still attached (#238).
  • The network address pin now takes effect (#255), the quarantine root the next artifact needs is kept (#246), and an attachment job fails cleanly when the disk does (#259).
  • The user is told when a code result lost data (#256), a clamped code duration stays an integer (#224), and shipped code raises real errors where it relied on assert (#222).
  • Engine hooks that receive host observations now declare them (#245).
  • "Allow once" no longer survives a restart, a locked execution store is no longer discarded as if it were corrupt, a store written by a newer version is set aside instead of stopping startup, and one bad cleanup record no longer turns off all retention (#287).

API and storage

  • Untrusted input gets a real status code instead of a 500 or a broken stream: a malformed Host header, an oversized image resize, a Last-Event-ID beyond SQLite's range, and ordinary states on the memory routes (#236).
  • Backup rotation and artifact retention no longer leave files behind on every startup and every artifact (#237).
  • Recovery discards the crashed database's write-ahead log instead of replaying it onto the restored backup (#230).
  • A corrupt memory file can be repaired from the app (#253).
  • The API accepts an IPv6 loopback host under Starlette 1.7's host parser (#272).
  • Database backups use SQLite's own backup interface, a failed backup no longer stops Cortex from starting, the chat backup is taken before a schema upgrade, and recovery keeps the crashed database's write-ahead log next to the quarantined file instead of deleting it, which replaces the behaviour described under #230 (#286).

Local models and the llama.cpp runtime

  • One corrupt GGUF file no longer hides every other model (#275; see "Corrupt Model Files" below).
  • Companion files are no longer offered as models, and the model list appears promptly (#270).
  • A GPU backend that cannot be fetched no longer blocks the CPU one (#232).
  • A startup timeout is reported instead of "starting" forever (#243).
  • A GGUF download cut short inside its tensor data is rejected (#244).
  • A one-second network hiccup no longer aborts the runtime download (#254).
  • The crash-loop warning no longer turns into its own opposite (#257).
  • The context window reported by Ollama is the one that is read (#249).
  • The llama.cpp child no longer inherits LLAMA_ARG_* overrides, runs a single slot, and has its web page switched off, and the context size the server actually loaded is read back (#289).

Launcher

  • A second Ctrl+C force-quits (#251).

Removed

  • The signed native worker that was never built (#215), the broker transport the coordinator never used (#217), the fake execution preview that shipped in the production app (#218), the cryptography dependency nothing imported (#219), and the consumed findings.json working file (#280).

Earlier entries

Long-form entries from before this layout, newest first. Their Version lines are pass names, not release numbers, unless a number is given. Entries dated after v1.0.0 (2026-01-20) have not been part of a published release; the 0.95.x entries describe Qt-era releases.

Stop Keeps the Answer

Date: 2026-09-25 Version: Bug-fix pass

  • Pressing Stop keeps the part of the answer you already watched appear. Since answers started streaming, Stop visibly erased text that was on screen: the turn was thrown away and the chat reloaded without it. The kept answer is saved with the chat, survives a restart, and is marked "Stopped" where the token count would be.
  • Only text that actually reached the screen is kept, so the answer never ends with words you did not see.
  • Stopping during translation keeps the untranslated answer instead of discarding a finished one.
  • Stopping a regenerated answer leaves the original in place, as before.
  • A stop that kept an answer no longer shows an error banner. When nothing had been written yet, the banner and its Retry remain, since Retry is then the way to ask again.

Cortex 2.0.0

Date: 2026-09-25 Version: 2.0.0

  • The declared version is 2.0.0, marking the rewrite as a different application from the Qt-era v1.0.0 so it sorts above it. There is no published download of 2.0.0; run Cortex from source.
  • The release workflow now builds the tag it was asked to build (a manual run used to build whatever main held), and only its upload step can write to the repository.

README Accuracy

Date: 2026-09-25 Version: Documentation pass

  • The README's pictures of the code feature now show code Cortex can actually run. They showed a program that imported a module and read a file off the disk, both of which Cortex refuses. The demo is now a loan calculation, and a test holds every staged demo task to the real validator and its real output.
  • All eight README screenshots are retaken from the current build, including the redesigned model picker and generation settings.
  • The README states plainly what the code feature can reach today: a scratch folder that is emptied for each run, no other programs, and a few public web requests. It no longer describes an image-transform feature that has no way to be started from the app, or a setup screen the app does not show.
  • Security reports now have somewhere to go: GitHub's private vulnerability reporting, which SECURITY.md pointed to, is switched on.

Corrupt Model Files

Date: 2026-09-25 Version: Bug-fix pass

  • One damaged .gguf file no longer hides all your other GGUF models. The fast model-file reader added on 2026-09-22 could crash on a file claiming an impossibly long entry. That error escaped the folder scan, and the model list quietly dropped every GGUF model, so your selected model showed as unavailable. A damaged file is now listed without its details, and every other model is listed normally.
  • If the GGUF scan fails anyway, Cortex now logs that it did, instead of showing an empty list without explanation.

Translation Failures

Date: 2026-09-25 Version: Bug-fix pass

  • A failed translation no longer breaks the answer while it streams in. Since the 2026-09-12 pass, a failed translation keeps the original answer, but the event announcing it was not one the app's own event list allowed, so the live stream broke at that point, and again on every reconnect. The answer only appeared once the app fell back to checking the job's status. A test now checks that every event the server can send is on that list.
  • The answer now says when it was left untranslated: "Couldn't translate this answer; showing the original." The note lasts until Cortex is closed, because the failure is not yet saved with the message.

GGUF Model Discovery

Date: 2026-09-22 Version: Model-scan pass

  • Models kept in subfolders are found. The scan looked in exactly one folder, which is not how any tool stores models -- each download lands in its own folder per repository -- so pointing Cortex at a models root showed only files loose at the top level, and pointing it at one model's folder showed that model alone. It now walks a bounded few levels, and a model's id carries its path.
  • Companion files are no longer offered as models. A multimodal projector and the later slices of a split model cannot be loaded on their own; selecting one used to fail much later as an unexplained restart loop.
  • The model list appears promptly. Reading the four details shown for each model used to parse the entire file including its tokenizer vocabulary, around eleven seconds per model; it now reads just the header block, so a folder of models lists in well under a second.

Real Token Streaming

Date: 2026-09-12 Version: Streaming pass

  • Responses now stream as the model writes them. Both local runtimes already produced a token stream; Cortex joined it, returned the finished answer, and only then replayed it to the interface in fixed slices, so a local model at a few tokens a second meant a spinner for the whole generation and then the answer all at once.
  • Reasoning streams into the reasoning pane on the same path, so a thinking model shows progress instead of silence.
  • Control blocks a reply may contain -- a memory proposal, a code-execution request, a legacy tag, or an inline reasoning trace -- are withheld from the live view, because the finished answer does not contain them. They no longer appear mid-answer and then vanish.
  • Title and translation calls are unchanged: they produce nothing to watch and still run as a single request.

Correctness Pass: Cancellation, Consent, Time, and Startup

Date: 2026-09-12 Version: Bug-fix pass

  • A cancelled code task can no longer run to completion. Only the terminal writes were guarded, so a worker that read the job just before Stop committed overwrote the cancellation and the program finished and reported success. Cancellation is now one-way in the execution store, which also makes the suite's one intermittent failure deterministic.
  • The brokered network capability refuses every non-global address. Python 3.13 reclassified 100.64.0.0/10 -- Tailscale's range, and common carrier-grade NAT -- as neither private nor reserved, so it had silently become reachable from an approved program.
  • Source too deeply nested for the parser is reported as invalid rather than escaping as an unhandled error, which had produced a 500 from the execution route and could interrupt a streaming turn.
  • A failed translation keeps the answer. The turn is persisted only after translation, so a failure there discarded a generation the user had already waited for; the untranslated answer is now returned with the reason beside it.
  • Ordinary writing turns no longer receive the code-execution contract. The admission gate matched prose such as "write a blog post about data science trends", which cost tokens and switched sampling to the coding profile, and its own explanation guard could never run.
  • Message timestamps record their UTC offset, so times no longer display shifted by the viewer's time zone and no longer jump when a chat reloads. Existing conversations are read back correctly without a migration, and forking a chat keeps each message's original time.
  • GGUF downloads are no longer capped at 8 GiB, which had refused most current mid-size models; free space remains the real limit.
  • Startup readiness probes ignore the system and environment proxy settings. A configured proxy cannot reach Cortex's own loopback socket, so Cortex failed to start on proxied machines with no explanation.
  • The launcher handoff secret survives a reload, so an expired session can still be re-exchanged instead of stranding the workspace on a retry that could never succeed.

Reliability, Performance, and Structure Pass

Date: 2026-09-04 Version: Quality pass

  • The settings database now runs in WAL mode. It had synchronous = NORMAL under a rollback journal, the one pairing SQLite does not make crash-safe. Backups checkpoint the write-ahead log before copying, and recovery drops sidecars belonging to the database it replaced.
  • range(0) is accepted by the code-execution validator instead of being rejected as unbounded, and the durable store and the resource governor now share one profile-name pattern rather than two that had drifted apart.
  • The execution event stream no longer runs synchronous SQLite on the event loop, and no longer closes after six seconds of silence -- which had been dropping the connection during the approval wait it exists to report on.
  • The transcript no longer re-renders on every streamed frame, history selection renders incrementally rather than rebuilding the whole transcript per stored message, and a turn no longer loads a full thread three times to read a title or a revision.
  • A worker that survives terminate and kill is reported rather than silently leaked.
  • Removed: unused memory scaffolding, the unused execution re-export barrel, the dead suggestions setting (existing workspaces still load), a vestigial second entry point, and the fake model gateway that shipped inside the packaged application.
  • One version string, read from cortex_backend.__version__, and a release workflow that refuses to build when a tag disagrees with it.
  • scripts/check.ps1 now runs the dependency lockfile check that CI enforces and verifies the installed environment against the declared pins.
  • The 53 API routes moved from a single 1,592-line function into one module per resource, with no change to the generated contract.

Local GGUF Models via a Managed llama.cpp Runtime

Date: 2026-08-07 Version: Local runtime pass

  • Cortex can serve a local .gguf file itself, through its own managed llama.cpp runtime, so a model no longer needs Ollama. GGUF models appear in the same picker as Ollama models, as gguf:<path>.
  • The runtime's llama-server is fetched once from ggml-org's llama.cpp releases and cached. Upstream publishes no checksums, so Cortex pins its own SHA-256 values in the source and verifies both the downloaded archive and every extracted file before launching it. Only the MIT-licensed Windows CPU and Vulkan builds are offered; the GPU backend setting is auto (try Vulkan, fall back to CPU), vulkan, or cpu.
  • A GGUF can be downloaded into the models folder by direct URL or Hugging Face repository, and each model's parameter size, quantization, and context length are read from the file itself.
  • The runtime's live status is shown beside the model picker.
  • The pinned llama.cpp build is moved forward by hand with tools/pin_llamacpp_release.py; CONTRIBUTING.md documents the procedure.

Cortex Native Web Shell & Self-Bootstrapping Package

Date: 2026-07-20 Version: Native web shell

  • Cortex now renders its React frontend in an owned pywebview/WebView2 desktop window instead of launching the user's installed browser.
  • The window uses a private Cortex-owned profile, remains tied to backend/Vite supervision, and closes the complete owned runtime when the native window exits.
  • Windows packaging is configured to verify and bundle Microsoft's signed Evergreen WebView2 bootstrapper, install the runtime only when absent, include the pywebview bridge, and emit a windowed executable without a companion console.
  • Second launches restore the existing Cortex window, while an explicit headless mode remains available for diagnostics and automation.

Cortex Web Modernization: Qt Removal and Release Candidate

Date: 2026-07-20 Version: Web modernization — Stage 7

  • The React/Vite web application and Python backend are now the only supported Cortex runtime; the obsolete desktop source, Qt dependencies, setup utility, and Visual Studio project artifacts were removed.
  • Legacy SQLite, JSON-chat, permanent-memory, and Windows registry settings remain readable. Legacy settings are imported through a Qt-free reader while the original source stays untouched for documented rollback.
  • Windows packaging includes prompt assets and the production frontend without requiring Node.js or a global Python installation at runtime.

Cortex Web Modernization: Launcher Cutover

Date: 2026-07-20 Version: Web modernization — Stage 6

  • The React/Vite web application is now the default python main.py runtime.
  • The Python launcher owns frontend builds, backend/Vite supervision, loopback readiness, authenticated browser handoff, single-instance behavior, and graceful UI shutdown.
  • Source builds use lock/source fingerprints and atomic bundle replacement; one-folder Windows packaging includes the frontend and does not require Node.js at runtime.
  • The temporary compatibility path used during migration was retired in the following release-candidate stage.

Cortex Web Preview: System Parity & Safe Migration

Date: 2026-07-20 Version: Staged web modernization — Stage 5

  • The preview web UI now exposes complete validated generation, translation, suggestion, memory, and model settings.
  • Legacy QSettings can be imported once into additive SQLite settings tables; migration status, invalid keys, and the verified database backup are exposed through diagnostics without writing back to the legacy source.
  • Permanent-memory edits and destructive clears are explicit, validated, and covered by browser and API tests.
  • Ollama connectivity, installed model tags, setup guidance, and streamed exact model-pull progress are available in the preview UI.

Cortex Remediation: Runtime, Storage, Safety, and Chat Reliability

Date: 2026-07-19 Version: Post-0.95.7 maintenance release

This staged maintenance series improves the reliability and maintainability of the desktop application without changing its Windows-first scope:

  • Runtime startup, connection checks, generation ownership, stale callbacks, and shutdown are now coordinated safely.
  • SQLite persistence and legacy migration are thread-safe, transactional, and recoverable; permanent memory writes use atomic replacement and backups.
  • Model-controlled memory actions are validated and destructive clears require confirmation. Rendered user, assistant, reasoning, and source content is sanitized, links are controlled, and sensitive model output is not logged.
  • Chat persistence state, fork indexes, regeneration, generated titles, context sizing, and inactive vector-memory initialization now follow explicit behavior.
  • Headless tests, bounded dependencies, Windows CI, and repository development instructions are included for ongoing maintenance.

Obsidian & Pumice: AI Memory Architecture

Date: 2025-10-16 Version: Critical Patch 0.95.7

This update addresses a series of critical, cascading failures in the AI's permanent memory system. The iterative process of teaching an AI nuanced rules is a philosophical and technical challenge. Previous attempts resulted in a system that oscillated between over-eagerly saving irrelevant data and being too paralyzed to save critical facts. This patch represents a stable, architectural solution, finally achieving the reliability and intelligence this feature was designed for.

The Overhaul: Achieving Reliable Memory

  • Fixed Critical Logic Failure: The most severe issue, where the AI could respond with a blank message and only a memory tag, has been eliminated. The system prompt was fundamentally re-architected to establish an unbreakable rule: a conversational response is always the primary goal, and memory functions are a silent, secondary task.
  • Intelligent Fact-Checking Heuristic: The AI is now equipped with a "Litmus Test" (Does this describe WHO the user is or just WHAT the user asked about?). This simple, powerful heuristic guides the model to correctly distinguish between a valuable user fact (like a name, profession, or stated interest) and a trivial conversational topic, resolving the root cause of irrelevant memories.
  • Balanced & Nuanced Instruction: The prompt has been carefully re-calibrated to be less punitive and more descriptive. This fixes the "prompt paralysis" that prevented the AI from saving legitimate user-stated facts, such as their name or personal interests, and ensures the system is neither over-eager nor over-cautious.
  • Robust & Consistent Behavior: By addressing these core architectural flaws in the prompt, the memory system is now significantly more robust, predictable, and useful. It can be trusted to build an accurate user profile over time without polluting the memory bank or failing its primary conversational duty.

Basalt & Slate: EULA Refinement & Architectural Decoupling

Date: 2025-10-16 Version: Feature Release 0.95.7

This update focuses on two distinct but philosophically linked areas: refining the user's first interaction with the application to be more deliberate and responsible, and re-architecting the AI's core identity to be more transparent and customizable.


Basalt: First-Run User Agreement

This update introduces a critical onboarding step to ensure user awareness and establish clear terms of use. My philosophy is that a responsible tool must be transparent about its nature. This feature formalizes the user's role and liability when interacting with local AI models.

The Feature: EULA on First Launch

  • One-Time Agreement: On the very first time the application is launched, a User Agreement & Liability dialog is now presented. This is a one-time event; once accepted, it will never appear again.
  • Clear Terms of Use: The dialog contains a standard End-User License Agreement (EULA) that outlines the user's responsibilities. It clarifies that the user, not the developer, is liable for the inputs provided to and the content generated by the local AI models they choose to run.
  • Mandatory Acceptance: The application will not proceed to the main interface until the user explicitly agrees to the terms by checking a box and clicking "Agree & Continue." This ensures a clear and deliberate acceptance of the terms.
  • Seamless & Themed Integration: The dialog is built using the same custom UI components as the rest of the application, ensuring it is fully themed (light/dark) and feels like an integrated part of the onboarding experience, not a generic system pop-up.
  • Persistent & Unobtrusive: The user's acceptance is saved permanently in the application's persistent settings (QSettings), guaranteeing a smooth, uninterrupted startup experience on all subsequent launches.

Refinement: Scroll-to-Agree & User Guidance

True agreement is not passive; it is an active and informed decision. The initial implementation of the EULA has been enhanced to ensure the user is genuinely presented with the full scope of the terms before being allowed to consent.

  • Engineered for Compliance: The "Agree & Continue" button is now intelligently disabled until the user has scrolled to the absolute bottom of the EULA text. This is a crucial step in due diligence, ensuring the entirety of the agreement has been made available.
  • Contextual Guidance: A disabled control without explanation is poor design. A tooltip has been added to the agreement checkbox, clearly instructing the user that they must scroll to the bottom to proceed. This removes ambiguity and friction from the onboarding process.
  • Robust & Intelligent Logic: The system is engineered to handle edge cases gracefully. If the EULA text is short enough to not require scrolling, the system recognizes this and enables the agreement option immediately.

Slate: Externalized AI Persona & Instructions

An application's logic should be distinct from its personality. This update fundamentally re-architects how the AI's core instructions are managed, moving them from being hardcoded within the application's source to residing in external, user-accessible text files. This is a critical step towards a more open, transparent, and customizable platform.

  • Decoupled Architecture: The AI's foundational instructions are no longer entangled with the Python code. They now live in two dedicated files: system_prompt.txt for the core identity and safety protocols, and memory_prompt.txt for the specific rules governing the permanent memory feature.
  • Empowering User Customization: This change transforms the AI's persona from a static element into a dynamic one. Advanced users can now directly edit these text files to tailor the AI's tone, behavior, and even its operational rules without ever touching a line of code.
  • Modular & Maintainable: Separating the core prompt from the memory prompt is a deliberate design choice. It makes the system's architecture cleaner and allows for the modular addition of future AI capabilities, each with its own externalized instruction set.
  • Performance-Aware Implementation: This externalization was engineered with performance in mind. The application reads these files only once upon first use and then caches them in memory, ensuring there is zero performance penalty on subsequent AI interactions.



Basalt: First-Run User Agreement

Date: 2025-10-15 Version: Feature Release 0.95.6

This update introduces a critical onboarding step to ensure user awareness and establish clear terms of use. My philosophy is that a responsible tool must be transparent about its nature. This feature formalizes the user's role and liability when interacting with local AI models.

The Feature: EULA on First Launch

  • One-Time Agreement: On the very first time the application is launched, a User Agreement & Liability dialog is now presented. This is a one-time event; once accepted, it will never appear again.
  • Clear Terms of Use: The dialog contains a standard End-User License Agreement (EULA) that outlines the user's responsibilities. It clarifies that the user, not the developer, is liable for the inputs provided to and the content generated by the local AI models they choose to run.
  • Mandatory Acceptance: The application will not proceed to the main interface until the user explicitly agrees to the terms by checking a box and clicking "Agree & Continue." This ensures a clear and deliberate acceptance of the terms.
  • Seamless & Themed Integration: The dialog is built using the same custom UI components as the rest of the application, ensuring it is fully themed (light/dark) and feels like an integrated part of the onboarding experience, not a generic system pop-up.
  • Persistent & Unobtrusive: The user's acceptance is saved permanently in the application's persistent settings (QSettings), guaranteeing a smooth, uninterrupted startup experience on all subsequent launches.



Version Update: 0.95.5 - 10/14/2025: Advanced Conversational Control

This update is centered on a single, powerful philosophy: a conversation should be a fluid, dynamic process that you, the user, can direct and refine at will. I've engineered two major new features—Regenerate and Fork—that transform the chat from a linear timeline into a branching tree of possibilities. These tools give you unprecedented control over the AI's output and the direction of your inquiry.

Pumice: AI Response Regeneration

You are no longer limited to the first answer the AI provides. With the new Regenerate feature, you can now prompt the model to rethink its last response, offering a new perspective, a different creative take, or a more refined solution. This is an essential tool for iterative creative work, technical problem-solving, and exploring the full potential of the model.

  • One-Click Reroll: A subtle "Regenerate" icon now appears below the most recent AI message. A single click will discard the previous response and generate a new one from your original prompt, keeping your workflow fast and intuitive.
  • Intelligent State Management: This feature is engineered with precision. The "Regenerate" button is only ever visible on the single, most recent AI response, preventing confusion and ensuring a clean, predictable user experience. The moment you send a new message, the option on the previous response disappears.
  • Seamless Integration: The regeneration process uses the same asynchronous, non-blocking architecture as a standard query, providing you with real-time status updates ("Analyzing," "Thinking...") while it formulates the new answer.

Geode: Conversational Forking

A single question can lead to a dozen interesting avenues of thought. The new Fork feature empowers you to explore them all without losing your place. You can now split any point in a conversation into a brand new, independent chat thread, preserving the context up to that moment.

  • Branch Your Conversation: A new "Fork" icon now appears below every AI-generated message. Clicking this button instantly creates a new chat, copying the entire history up to that point. This allows you to pursue a tangent or explore an alternative line of questioning in a clean, separate workspace.
  • Intelligent & Automatic Titling: Forked chats are named with clarity and precision. The system automatically takes the original chat's title and appends a "Thread" counter (e.g., "Python Scripting" becomes "Python Scripting Thread:2"). If you fork from a thread that is already a fork, it intelligently increments the number ("Python Scripting Thread:2" becomes "Python Scripting Thread:3"), maintaining a clear and logical hierarchy.
  • Atomic & Instantaneous Creation: The underlying database logic has been engineered for performance and reliability. The creation of a forked chat is an atomic transaction, meaning the new thread and all its copied messages are saved to the database in a single, instantaneous, and failure-proof operation.



Version Update: 0.95.0 - 10/14/2025: Code Rendering & UI Polish

This update introduces a completely re-architected system for displaying code within conversations. My philosophy is that the tools you use should be as well-crafted as the work you create with them, and this feature elevates code from a simple text element to a rich, interactive, and professional component of the UI.

Onyx: Professional Code Rendering & Syntax Highlighting

  • Dedicated Code Container: Code blocks now live in a dedicated, professionally styled container that is visually distinct from conversational text. This container features a clean header that displays the detected programming language, providing immediate context.
  • One-Click Copy Functionality: Each code block features a "Copy" button in its header, allowing you to extract snippets with a single click. The button provides subtle "Copied!" feedback, streamlining your workflow and eliminating the need for manual selection.
  • Native, Theme-Aware Highlighting: The rendering pipeline uses a high-performance, Qt-native QSyntaxHighlighter. This ensures that syntax highlighting is not only fast and accurate but also seamlessly adapts to both light and dark themes for perfect readability.
  • Intelligent Scrolling & Flawless Geometry: The container is independently scrollable for longer code blocks, preserving a clean chat layout. The widget has been meticulously engineered to display flawless, rounded corners that are pixel-perfect in all themes.



Granite: Critical UI Interaction Fix

Date: 2025-10-14 Version: Hotfix 0.94.4

This hotfix addresses a critical bug that made the "Copy" and "Copy All" actions for chat messages non-functional.

The Issue

The "Copy" (for selected text) and "Copy All" actions in the context menu were failing, either doing nothing or copying a blank string to the clipboard. This broke a fundamental user interaction.

Root Cause

The investigation revealed two underlying problems:

  1. State Loss on Focus Change: When the context menu appeared, it took focus from the message text. This caused the operating system to clear any active text selection before the copy action could read it.
  2. Incorrect Data Handling: The widget was not reliably storing the original, plain-text version of the message, causing "Copy All" to fail.

The Solution

The event-handling logic has been re-engineered for stability:

  • Proactive Text Capture: The application now captures any selected text the instant a right-click occurs, before the selection is cleared by the menu appearing.
  • Robust Data Management: The logic has been corrected to ensure the full, unformatted message is always available for the "Copy All" command.

Impact on Users

All clipboard functions within the chat view are now fully restored and reliable. You can copy selected text snippets and full messages without issue.




Version Update - 10/13/2025: Update Notification System & Stability Fixes

This update introduces a new, non-intrusive update notification system to keep you informed of the latest features and fixes, alongside a crucial stability enhancement for the update-checking process itself.

Alabaster: Smart Update Notifications

Staying up-to-date should be effortless. This release introduces a new, intelligent update notification system designed to be informative without being disruptive. My philosophy is that you should be in control of your workspace, and this feature reflects that.

  • Automatic & Asynchronous Checking: When the application starts, it now performs a silent, one-time check in the background to see if a new version is available. This process is fully asynchronous, meaning it will never freeze or slow down your startup experience.
  • Non-Intrusive Notification: You won't be interrupted by pop-ups. Instead, if a new version is detected, a clean and clear notification will be waiting for you within the Settings dialog. This allows you to check for updates on your own terms.
  • Clear Status, Always: The Settings dialog now provides transparent feedback on the update check's status. You will always know if the application is up-to-date, if an update is available, or if the check encountered a network error.
  • Robust Cache-Busting: The web request has been engineered to be highly robust. It sends specific headers that instruct servers and proxies to bypass their caches, ensuring the check always fetches the latest, live version information and never gives a false negative due to stale, cached data.


Update Log: Critical Hotfix for AI Reasoning (Chain-of-Thought) Display

Date: 2025-10-13 Version: Hotfix 0.94.3

Summary (TL;DR)

A critical hotfix has been deployed to address a bug where the AI's step-by-step reasoning was not being displayed in the chat interface. The "View Reasoning" button, which reveals the model's thought process, is now fully functional again for all compatible models.

Expected .exe push update soon


The Issue: Missing "View Reasoning" Button

We identified a critical issue reported by our users where the "View Reasoning" button was consistently absent from the AI assistant's chat bubbles. This feature provides transparency into the model's Chain-of-Thought (CoT) process, allowing users to understand how an answer was formulated. Its absence was a significant regression in functionality and user experience.

Root Cause Analysis

After a thorough investigation, we determined the root cause was an upstream change in the Ollama API's response structure (specifically in versions 0.12.5 and newer).

In a welcome move to improve clarity, the Ollama API now separates the model's reasoning process into a dedicated thinking field within the response. Previously, this reasoning block was embedded directly inside the main content field.

Our application's data parsing logic was still programmed to look for the reasoning in the old, embedded location. When the new API structure was encountered, our system failed to find the thinking data, and as a result, the UI correctly determined there was no reasoning to display.

The Solution: A Robust and Backward-Compatible Fix

The SynthesisAgent, responsible for handling communication with the LLM, has been updated with smarter response-handling logic:

  1. Primary Parsing Path: The system now correctly looks for and extracts the reasoning from the new, dedicated thinking field provided by modern Ollama versions.
  2. Fallback Mechanism: To ensure full compatibility, we have retained the old parsing logic as a fallback. If the thinking field is not present in a response, the system will then scan the main content for the inline reasoning block.

This dual approach ensures that the "View Reasoning" feature works seamlessly for users running any version of the Ollama service, providing both forward compatibility with the latest updates and backward compatibility for those on older installations.

Impact on Users

With this fix deployed, the "View Reasoning" functionality is fully restored. You can once again gain valuable insight into the AI's problem-solving process. No action is required on your part.

We extend our sincere thanks to the community member who provided the detailed logs and API outputs that allowed for a swift diagnosis and resolution of this issue. Your feedback is invaluable in helping us maintain the quality and reliability of the application.




Version Update - 10/12/2025: Architectural Overhaul & Major New Features

This is a landmark update that touches nearly every part of the application. I've focused on rebuilding core systems for performance and reliability, while also introducing powerful new features and quality-of-life improvements based on how I see the app evolving.

Bedrock: Architectural Overhaul - Next-Generation Data Storage

I've fundamentally rebuilt how conversational data is stored and managed, transitioning from a system of individual text files to a robust, high-performance SQLite database. This architectural shift moves the application to a professional-grade storage solution, delivering massive improvements in speed, reliability, and scalability.

  • Enhanced Performance & Scalability: You will experience a dramatic increase in speed, especially when loading your chat history. What once required scanning multiple files on disk is now an instantaneous, indexed database query. The application will feel significantly more responsive, regardless of whether you have ten chats or ten thousand.
  • Rock-Solid Data Integrity: The new database system is transactional, which protects your chat history against corruption from unexpected application crashes or power failures. This ensures your conversations are always saved safely.
  • Seamless, Automatic Migration: For existing users, this transition is completely effortless. On its first run after the update, the application will automatically detect your old chat files and migrate them into the new database. All of your history will be preserved with no action required from you.
  • Foundation for Future Features: This new architecture isn't just about improving what's here; it's about unlocking what's next. It lays the groundwork for powerful future capabilities, such as a full-text search across all of your past conversations.

Flint: Automated Installer & Onboarding

Getting started previously required manually finding, downloading, and installing Ollama, followed by using command-line tools to pull AI models. This process could be a significant barrier.

The new Corted Startup utility completely transforms this experience into a simple, guided workflow. My goal is to remove technical barriers and empower every user to get up and running in minutes.

  • Guided Ollama Installation: The utility now presents direct download links for Windows and macOS and a one-click copy command for Linux. There's no more need to search for installation instructions.
  • Integrated Model Manager: Forget the command line. You can now browse a curated list of available AI models directly within the application. Select the model you want, click "Pull," and monitor the download and installation progress in real-time.
  • A Polished, All-in-One Interface: This entire process is wrapped in a clean, modern interface featuring a draggable window and a light/dark theme toggle to match your workspace. It provides a professional and centralized starting point for the entire Corted experience.

Quartz: AI Customization - User-Defined System Instructions

This update introduces a powerful new dimension of control over the AI assistant. You can now define a persistent persona and set global behavioral rules through the new System Instructions feature, allowing you to tailor the AI's core identity to your specific needs.

This feature was engineered with extreme care. The underlying AI prompt has been restructured to create an unmistakable hierarchy, ensuring the model perfectly understands its core function, your custom instructions, and the context of your conversation.

  • Set a Persistent Persona: Accessed via the Settings menu, the new System Instructions dialog allows you to provide high-level directives that apply to every conversation.
  • Precision Control Over AI Behavior: Your custom instructions are given the highest priority, giving you unprecedented influence over the tone, style, and structure of the AI's responses.
  • Robust & Unambiguous Prompting: The AI's core prompt has been meticulously re-architected to ensure your instructions are integrated without conflicting with its primary system functions, leading to a more predictable and reliable output.

Obsidian: Advanced Model Controls

Building upon the foundation of AI customization, this update provides granular control over the core parameters of the language model itself. These advanced settings have been moved to their own dedicated dialog, ensuring a clean, uncluttered interface that gives power users the tools they need to fine-tune the AI's performance.

  • Dedicated Control Panel: A new "Advanced Model Settings" dialog provides a focused workspace for adjusting the AI's inner workings without cluttering the main settings panel.
  • Temperature Control: An intuitive slider allows you to precisely manage the AI's creativity. Lower the temperature for more deterministic, factual responses, or raise it to encourage more novel and imaginative output.
  • Context Window Management: Directly set the size of the model's conversational memory (context window). Increase it for longer-term context retention in complex conversations, or decrease it to optimize performance and memory usage.
  • Reproducible Outputs with Seeding: A seed value can now be set to ensure the AI produces the exact same response to the same prompt every time, a crucial feature for testing, development, and content generation. A "Random" button is provided for convenience.

Amethyst: Keyboard Shortcuts - Command at the Speed of Thought

A truly powerful tool should feel like an extension of your own thoughts, minimizing the friction between intent and action. This update introduces a comprehensive suite of keyboard shortcuts, designed to keep your hands on the keyboard and your mind focused on the conversation. My goal is to elevate the application from a simple point-and-click interface to a high-velocity command center for power users.

  • New Conversation (Ctrl+N): Instantly start a fresh chat without reaching for the mouse, keeping your workflow seamless.
  • Access Settings (Ctrl+,): Quickly open the settings dialog with a standard, universal shortcut to tweak the AI's model or appearance on the fly.
  • Focus Input (Ctrl+L): Immediately jump to the chat input field from anywhere in the application. This is a massive time-saver for rapid-fire questioning.
  • Close Window (Ctrl+W): A standard, convenient way to close the application window when your session is complete.
  • Cross-Platform Native Feel: These shortcuts are intelligently mapped, automatically translating to Cmd on macOS to ensure a native and intuitive experience on any operating system.

Diamond: Hierarchical UI & Blur Overhaul

I've re-architected the application's dialog system to create a proper sense of depth and focus. The previous implementation could incorrectly blur parent dialogs, leading to a confusing user experience. This has been resolved with a more intelligent, stack-based management system.

  • Correct Visual Hierarchy: The application now correctly tracks the layering of all open windows. When a new dialog appears, only the windows behind it are blurred, ensuring the active window is always sharp and clearly in focus.
  • Robust Dialog Stacking: The new system can flawlessly handle any number of nested dialogs (e.g., opening "Manage Memories" from within the "Settings" dialog) while maintaining the correct visual state. This is a crucial fix for UI professionalism and usability.

Sandstone & Keystone: UX and Logic Enhancements

These updates focus on improving the core conversational experience and providing greater control over your data.

  • New Feature: Clear All Chat History (Sandstone): You now have the ability to permanently delete your entire chat history. This action performs a "clean sweep" of all recorded conversations, resetting your history to a blank slate. This option is accessible by right-clicking the "+ New Chat" button. A confirmation prompt will appear to prevent accidental data loss.
  • Resolved Repetition Bug (Keystone): I've fixed a core logical issue where the AI would sometimes perceive the user's first message as a repeated statement, causing odd conversational artifacts (e.g., greeting you "again"). This is now fully resolved, significantly improving the model's contextual understanding from the very first turn.

Shale: UI Polish & Theming Fixes

A powerful tool should also be a pleasure to use. This update refines the application's user interface by addressing several visual inconsistencies and bugs.

  • Instantaneous Theme Switching: All UI elements now update their appearance instantly when switching between light and dark modes.
  • Improved UI Clarity: Corrected an issue where the text field in the System Instructions dialog could blend into the background in light mode.
  • Stylesheet Stability: Fixed a minor bug that could cause stylesheet parsing errors, leading to more stable and robust rendering of all UI components.