Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,8 @@

## Unreleased

- Unreleased (on main, not in v0.20.0): on SQLite, `sources delete`, `sources prune --superseded` and `import-markdown --supersede` now scrub an open loop that names the source only in its metadata: the id, or `source:<id>`, as the text under `source_id`, `source_ids`, `source_ref`, `source_refs`, `source_references` or `selected_source_ids` at any depth. That is the rule the open-loop lookup of a source already used. They blanked only loops whose `source_id` column held the id, so a loop with an empty column kept its title, description and metadata, and they reached a new export. The delete preview now counts the same loops as the receipt. The rule reads the id only as stored or as `source:<id>`: a loop that names the source in any spelling other than those two keeps its text, although the memory rule reads several other spellings. No migration is required.
- Unreleased (on main, not in v0.20.0): on SQLite, `sources delete`, `sources prune --superseded` and `import-markdown --supersede` blank an open loop that names the source in any spelling the saved-quote reader names, including capitals, no hyphens, braces, `urn:uuid:`, a list and JSON text. The lookup, the preview and the reader that withholds those ids use the same spellings, including JSON escapes and repeated keys. Lookup and removal both decode JSON text under source and memory references and inside trace fields. Admitted references and surrounding values stay, and the owner receives live references unchanged. One pass over the user's loops answers for every source of the command. v0.20.0 blanked only the id as stored or `source:<id>`, so any other spelling kept the loop's text. No migration is required.
- Unreleased (on main, not in v0.20.0): on SQLite, `sources delete`, `sources prune --superseded` and `import-markdown --supersede` now scrub an open loop that names the source only in its metadata: the id, or `source:<id>`, as the text under `source_id`, `source_ids`, `source_ref`, `source_refs`, `source_references` or `selected_source_ids` at any depth. That is the rule the open-loop lookup of a source already used. They blanked only loops whose `source_id` column held the id, so a loop with an empty column kept its title, description and metadata, and they reached a new export. The delete preview now counts the same loops as the receipt. Correction (2026-10-05): those other spellings are blanked too, by the same reader the saved-quote fence uses. No migration is required.
- Unreleased (on main, not in v0.20.0): the SQLite derived-label repair now updates and records each row under the id it is stored with (SQLite keeps capitals, missing hyphens, braces and `urn:uuid:` as written) and matches rows by the normalised id only to read the graph. If two stored spellings of one id exist, each is compared with the label its own recorded inputs give it, so neither is left at its old label while the other is relabelled, and neither is lowered. An update that changes no row stops the repair with `DerivedDomainRepairError` before any event or the completion stamp is written, on open and on restore; migration `20261004_0095` refuses the same way. Before, a derived memory stored under such an id kept its label while its relabel event and the stamp were written, and the pair survived export and import. A PostgreSQL uuid column was not affected. A vault that the earlier repair already stamped keeps such a row at its old label until a restore repairs it again. No migration is required.
- Unreleased (on main, not in v0.20.0): v0.20.0 could store a consolidation report with a sensitivity below the memories it prints. The report took its sensitivity from the near-duplicate clusters only, so a run whose proposals were roll-up cards had no cluster to take it from and was stored as `unknown`. A key whose ceiling is below those memories (a `trusted_local_agent` key for confidential ones, a `read_only_agent` key for private ones) could then read the card topic and the member ids through `GET /v0/vnext/artifacts/{id}`. The report now takes its domain and its sensitivity over every row it names: the cluster members, the members of every proposed roll-up group, the members of the groups that a skip line names by key (`topic:...`, `entity:...`, `semantic:cluster-<id>`), the pending, accepted, expired or held roll-up cards it names by id, and the sources that the `source_refs` of its cluster members name, which the report still prints. The open-loop review is labelled over the sources whose ids it prints as well as over its loops. The run digest of both reports covers those sources, so a source that was reclassified makes a new report instead of returning the earlier one. The first run after the upgrade over loops that link a source, or over cluster members that cite one, makes one new report, and a run that names no source keeps the digest it had. A report whose inputs are all unrestricted is now stored with the label of those inputs, `internal` where it was `unknown`, and every permission profile reads the two the same way. Reports stored by v0.20.0 keep their labels, because the stored-row repair does not change sensitivity. The other report producers (daily brief, weekly synthesis, connection and contradiction reports, project updates, staleness reports) were checked and already label over every row they print. Memories and sources are counted apart when the label is taken, because an id is unique only within its own table: a source that shares a memory's id can no longer replace that memory's label. No migration is required.
- Unreleased (on main, not in v0.20.0): v0.20.0 `alice-memory export` copied the vault into a private folder in the temporary directory and removed it only on a normal exit or an exception, so a SIGKILL or a power loss left a plaintext copy of the vault there until the system cleaned the directory. `sources list`, the `sources delete` and `sources prune` previews and `import-markdown --dry-run` on main use the same kind of folder, and `alice-memory import` in v0.20.0 does the same for a copy of the file it imports. Each folder now holds a small marker that names the process that made it (its PID and start time), and every command that makes one first removes the folders of processes that are gone, or whose PID now belongs to a process that started at a different time. The sweep only touches a real directory (never a symlink) directly under the temporary directory, with one of the three snapshot name prefixes, owned by the current user, mode 0700, holding a valid marker. It leaves everything else alone, keeps a folder whenever it cannot tell whether the process is running, and never fails the command. The `import-markdown --dry-run` copy of the vault and of `sleep_proposals.jsonl` is now created with mode 0600. No migration is required.
Expand Down
18 changes: 11 additions & 7 deletions apps/api/src/alicebot_api/source_commands.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
from alicebot_api.vault_sleep import SleepError
from alicebot_api.vnext_stores.memory_lifecycle_common import is_redacted_memory
from alicebot_api.vnext_stores.sqlite.source_retirement import (
CandidateScrubRefused, citing_memories_by_source, optimize_scrub_indexes, source_open_loop_count)
CandidateScrubRefused, citing_memories_by_source, open_loops_naming_sources, optimize_scrub_indexes)

RETAINED_DATA = (
"Source and import events keep prior titles, hashes and import folder paths. "
Expand Down Expand Up @@ -40,8 +40,10 @@ def _targets(store, args):


def _preview(store, rows):
# One pass over the memories answers for every source of the preview.
citing=citing_memories_by_source(store,[str(row['id']) for row in rows])
# One pass over the memories, and one pass over the loops, answer for every source of the preview.
source_ids=[str(row['id']) for row in rows]
citing=citing_memories_by_source(store,source_ids)
loops=open_loops_naming_sources(store,source_ids)
result = []
for row in rows:
sid=str(row['id'])
Expand All @@ -51,8 +53,7 @@ def _preview(store, rows):
"(SELECT 1 FROM source_chunks c WHERE c.user_id=p.user_id AND c.id=p.source_chunk_id "
"AND c.source_id=?))) AS provenance_quotes",
(store.user_id,sid,store.user_id,sid,sid))
# The loops the scrub blanks: the one rule of the source reverse lookup, not the column alone.
counts['open_loops']=source_open_loop_count(store,sid)
counts['open_loops']=len(loops[sid])
memories=[memory for memory in citing[sid] if not is_redacted_memory(memory)]
counts['candidate_memories']=sum(memory['status'] in {'candidate','needs_review','rejected'} for memory in memories)
counts['memories_citing_replaced']=[str(memory['id']) for memory in memories
Expand Down Expand Up @@ -91,9 +92,12 @@ def run_sources(args):
with store.savepoint():
rows=_targets(store,args)
# One pass over the memories finds the citing memories of every source; each scrub reads those again.
citing=citing_memories_by_source(store,[str(row['id']) for row in rows])
source_ids=[str(row['id']) for row in rows]
citing=citing_memories_by_source(store,source_ids)
loops=open_loops_naming_sources(store,source_ids)
receipts=[{'id':str(row['id']),**store.scrub_source(str(row['id']), optimize=False,
citing_ids=[str(memory['id']) for memory in citing[str(row['id'])]])} for row in rows]
citing_ids=[str(memory['id']) for memory in citing[str(row['id'])]],
loop_ids=[str(loop['id']) for loop in loops[str(row['id'])]])} for row in rows]
if receipts:
optimize_scrub_indexes(store)
print(json.dumps({'deleted_count':len(receipts),'deleted':receipts,'retained_data':RETAINED_DATA},sort_keys=True))
Expand Down
11 changes: 10 additions & 1 deletion apps/api/src/alicebot_api/vnext_capture.py
Original file line number Diff line number Diff line change
Expand Up @@ -1346,11 +1346,14 @@ def capture_source(self, source_input: SourceCaptureInput) -> CaptureResult:
return replace(result, kept_reason="matches_other_live_source")
retired = []
citing = []
cached = getattr(self, "_open_loop_names", None)
for row in matches:
if str(row['id']) == str(result.source_id):
continue
loop_ids = None if cached is None else [str(loop["id"]) for loop in cached.get(str(row["id"]), [])]
counts = getattr(self.store, "supersede_source")(str(row['id']), superseded_by=result.source_id,
allow_looser_classification=policy.allow_looser_classification, dry_run=policy.dry_run)
allow_looser_classification=policy.allow_looser_classification, dry_run=policy.dry_run,
loop_ids=loop_ids)
retired.append({"id": str(row['id']), "title": printed_source_label(row.get('title'))})
citing.extend(counts['memories_citing_replaced'])
return replace(result, superseded=tuple(retired), memories_citing_replaced=tuple(dict.fromkeys(citing)))
Expand Down Expand Up @@ -1845,6 +1848,11 @@ class PreviewRollback(Exception):
scan = getattr(self.store, "markdown_sources_by_path", None)
with self.store.savepoint() if callable(scan) or dry_run else nullcontext():
self._markdown_path_index = scan() if callable(scan) else {}
self._open_loop_names = None
if policy.mode != "off":
from alicebot_api.vnext_stores.sqlite.source_retirement import open_loops_naming_sources
indexed = [str(row["id"]) for rows in self._markdown_path_index.values() for row in rows]
self._open_loop_names = open_loops_naming_sources(self.store, indexed)
result = self._import_markdown_folder(folder, domain=domain, sensitivity=sensitivity,
max_file_bytes=max_file_bytes, policy=policy)
if dry_run:
Expand All @@ -1853,6 +1861,7 @@ class PreviewRollback(Exception):
return replace(result, dry_run=True)
finally:
self._markdown_path_index = None
self._open_loop_names = None
return result

def _import_markdown_folder(
Expand Down
18 changes: 17 additions & 1 deletion apps/api/src/alicebot_api/vnext_open_loop_references.py
Original file line number Diff line number Diff line change
Expand Up @@ -67,11 +67,12 @@
from __future__ import annotations

import inspect
import json
import re
from collections.abc import Callable, Iterable, Mapping, Sequence
from uuid import UUID

from alicebot_api.vnext_source_fence import SOURCE_REFERENCE_KEYS, SourceReadFence, cited_source_ids
from alicebot_api.vnext_source_fence import SOURCE_REFERENCE_KEYS, SourceReadFence, _json_container, cited_source_ids

JsonObject = dict[str, object]

Expand Down Expand Up @@ -131,6 +132,9 @@ def withhold_unreadable_references(
if memory is not None:
memory_ids.add(memory)
referenced.add(memory)
named = cited_source_ids(row.get("metadata_json")).named
metadata_ids.update(named)
referenced.update(named)
_collect_ids(row.get("metadata_json"), metadata_ids, referenced, at_reference=False, depth=0)
source_rows = _rows_by_id(store, sorted(source_ids | metadata_ids), bulk="get_sources_by_ids", single="get_source")
memory_rows = _rows_by_id(store, sorted(memory_ids | metadata_ids), bulk="get_memories_by_ids", single="get_memory")
Expand Down Expand Up @@ -276,7 +280,13 @@ def _collect_ids(value: object, found: set[str], referenced: set[str], *, at_ref
if depth > _METADATA_MAX_DEPTH:
return
if isinstance(value, str):
decoded = _json_container(value)
if decoded is not None:
_collect_ids(decoded, found, referenced, at_reference=at_reference, depth=depth + 1)
return
ids = _ids_in_text(value)
if at_reference:
ids |= set(cited_source_ids(value).named)
found.update(ids)
if at_reference:
referenced.update(ids)
Expand Down Expand Up @@ -364,6 +374,12 @@ def _scrub(value: object, *, depth: int, withheld: frozenset[str]) -> object:
if depth > _METADATA_MAX_DEPTH:
return _DROPPED
if isinstance(value, str):
decoded = _json_container(value)
if decoded is not None:
checked = _scrub(decoded, depth=depth + 1, withheld=withheld)
if checked is _DROPPED:
return _DROPPED
return value if checked == decoded else json.dumps(checked)
return _scrub_text(value, withheld=withheld)
if isinstance(value, Mapping):
output: dict[object, object] = {}
Expand Down
17 changes: 3 additions & 14 deletions apps/api/src/alicebot_api/vnext_stores/sqlite/graph_open_loops.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@
GRAPH_EDGE_COLUMNS,
OPEN_LOOP_COLUMNS,
)
from alicebot_api.vnext_stores.sqlite.open_loop_source_reference import open_loop_source_reference_sql
from alicebot_api.vnext_stores.sqlite.source_retirement import open_loops_naming_sources
from alicebot_api.vnext_stores.sqlite.primitives import (
_iso_or_none,
_iso_or_now,
Expand Down Expand Up @@ -664,22 +664,11 @@ def find_open_loop_by_automation_digest(
)

def list_open_loops_referencing_source(self, *, source_id: str, limit: int = 500) -> list[VNextRow]:
"""Bound open loops related to one source before LIMIT."""
"""Open loops that name one source, by the spellings ``cited_source_ids`` names, before LIMIT."""

if limit < 1:
raise ValueError("limit must be positive")
reference, reference_params = open_loop_source_reference_sql(source_id)
return self._fetch_all(
f"""
SELECT {", ".join(OPEN_LOOP_COLUMNS)}
FROM open_loops
WHERE user_id = ?
AND {reference}
ORDER BY updated_at DESC, created_at DESC, id DESC
LIMIT ?
""",
(self.user_id, *reference_params, limit),
)
return open_loops_naming_sources(self, [source_id]).get(str(source_id), [])[:limit]

def list_open_loops(
self,
Expand Down
Loading
Loading