Skip to content

Commit a5aa5b5

Browse files
committed
Fix: Move cache spec to separate file.
1 parent ee0d258 commit a5aa5b5

3 files changed

Lines changed: 290 additions & 272 deletions

File tree

Lines changed: 285 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,285 @@
1+
Cache spec
2+
==========
3+
4+
The cache stores the location (file name and ID) and contents (inner HTML or
5+
code/doc blocks for fragments) of a target. Targets are HTML elements (excluding
6+
`<fragment>`s) with an ID.
7+
8+
The goal of the cache is to support cross-references and gather elements, and to
9+
ensure that all IDs are unique within a project. This means that
10+
cross-references and gather elements persist across moving or renaming files,
11+
since the IDs will be found in the cache.
12+
13+
Non-project files support a subset of this functionality: the "project" consists
14+
only of the current file. Only targets, gather elements, cross-references, and
15+
fragments to items within the file work as expected; references to other files
16+
do not.
17+
18+
The cache reflects data read directly from disk/IDE; content edited in the
19+
Client does not update the cache until it's written to disk/IDE, at which point
20+
the cache much re-process this unknown file. Since the Client is designed around
21+
an autosave principle which updates disk/IDE regularly, there's little gap
22+
between the two.
23+
24+
The cache hydrates data to the Client; the dehydration routes are responsible
25+
for removing all cache hydration artifacts.
26+
27+
Cross references
28+
----------------
29+
30+
A `<xref ref="id"></xref>` is a cross reference to a `Target` or a gather
31+
element. The `id` specifies the destination; the cache then hydrates the
32+
contents based on the location and contents of the target of the provided `id`
33+
to e.g. `<xref ref="id" contenteditable="false"><a
34+
href="../path/to/page#id">Inner HTML from target</a></xref>`; this hydrated form
35+
is only present in the Client.
36+
37+
Details:
38+
39+
* This element does not allow an `id` attribute.
40+
41+
* The inner HTML is always taken before cache hydration, to prevent circular
42+
dependencies:
43+
44+
```html
45+
<h1 id="a">See <xref ref="b"></xref></h1>
46+
<!-- file a -->
47+
<h1 id="b">See <xref ref="a"></xref></h1>
48+
<!-- file b -->
49+
```
50+
51+
Including cache hydration would cause the inner HTML to be updated each time
52+
the file is processed, outdating the other file.
53+
54+
* If the `id` referred to isn't found or refers to a duplicate id, the inner
55+
text is instead an appropriate error message.
56+
57+
* If the cross reference is to a gather element, the text is the gather
58+
element's inner HTML, not the gathered code/doc blocks.
59+
60+
Fragments and gather elements
61+
-----------------------------
62+
63+
A gather element such as `<h3 id="bar" data-gather="id1 id2...">Bazzy
64+
things</h3>` is a `Target` with the `data-gather` attribute. It becomes a list
65+
of the contents of fragments it refers to after hydration by the cache. An
66+
example fragment tag, after cache hydration: `<fragment id="id1"
67+
contenteditable="false">See <a href="path/to/gather#bar">Bazzy things</a>, <a
68+
href="path/to/another/gather#zap">Zappy things</a></fragment>`. A fragment's
69+
content by default includes the contents of the current doc block and the
70+
contents of the following code/doc block; fragments are not allowed in Markdown
71+
documents (in the case, the fragment contents consist of an error message).
72+
Fragments may include the `following` attribute to enclose a specific number of
73+
the following code/doc blocks; for example, `<fragment id="bar"
74+
following="3"></fragment>` includes the current doc block along with the next 3
75+
code/doc blocks; `following` must be a whole number.
76+
77+
Details:
78+
79+
* Fragment contents may not include a gather element; in this case, the gather
80+
element list of contents will simply include an error message.
81+
* Fragments do support indirection: gather element A includes contents from
82+
fragment B, which contains a cross reference to target C. Changes to target C
83+
makes B and A outdated.
84+
* If a gather element refers to an `id` that is a `Target` or a `GatherElement`,
85+
not a `Fragment`, the resulting output for this in the list of fragments is an
86+
error message.
87+
* If a gather element refers to an `id` that wasn't found or is a duplicate, its
88+
contents will be replaced by an error message.
89+
* Fragments store an HTML rendering of the code and doc blocks they contain,
90+
excluding the content produced by hydrating the `<fragment>` tags, to avoid
91+
duplication and circular dependencies. See layer 4 under `Design`: this
92+
exclusion is what keeps a gather element and the fragments it lists from
93+
outdating each other forever. The HTML rendering of a fragment reproduces the
94+
layout of the source it came from: each doc block includes its indent, each
95+
line of a code block is preceded by that line's number, and the two are
96+
aligned -- a doc block and a line of code indented equally in the source begin
97+
in the same column, with the line numbers in a gutter of their own to the left
98+
of both.
99+
* The backlinks a `<fragment>` hydrates to are derived from the gather elements
100+
which list it: for each such element, its containing file, its `id`, and its
101+
inner HTML (the link text). A fragment's rendered output therefore depends on
102+
the gather elements which reference it -- the reverse of the direction in
103+
which references are written; see `Design`.
104+
* A `<fragment following=0>` is valid; it contains only the current doc block.
105+
* The `following` attribute is clamped if it would exceed the number of code/doc
106+
blocks in the document.
107+
* If the `following` value cannot be parsed to a whole number, an error message
108+
replaces the fragment content.
109+
* A gather element, as a type of `Target`, requires an `id`. A `data-gather`
110+
attribute on an element without an `id` produces an error message which
111+
requests the missing `id`.
112+
* All ids must be valid
113+
[CSS identifiers](https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/es/ident)
114+
per
115+
[MDN recommendations](https://developer.mozilla.org/en-US/docs/Web/HTML/rence/Global_attributes/id).
116+
Invalid ids produce error messages in the hydrated tag content.
117+
118+
Example hydration of the gather tag `<h3 id="bar" data-gather="id1 id2...">Bazzy
119+
things</h3>`:
120+
121+
```html
122+
<h3 class="cc-gather" id="bar" data-gather="id1 id2...">Bazzy things</h3>
123+
<div class="cc-gather-items" contenteditable="false">
124+
<p class="cc-gather-item-link">
125+
From <a href="link/to/first/tag#id1">Path to file</a>:
126+
</p>
127+
(first item content) ...
128+
<p class="cc-gather-item-link">
129+
From <a href="link/to/last/tag#idn">Path to file</a>:
130+
</p>
131+
(last item content)
132+
</div>
133+
```
134+
135+
Search
136+
------
137+
138+
The cache supports searching the (cleaned) inner HTML of all `Target`s and
139+
gather elements; search does not include `Fragment` contents.
140+
141+
### Auto-assignment of ids
142+
143+
If `id="*"` on either a fragment, target, or gather element, the cache replaces
144+
this with an random autogenerated `id` placed in the resulting HTML contents,
145+
but this new `id` is not yet recorded in the cache. The file is then marked as
146+
`Unknown` when it is saved; when re-read, this `id` is then incorporated into
147+
the cache. This helps avoid cases where the cache and file contents become
148+
unsynchronized: if the `id` is placed in the cache before the write and the
149+
write fails, or if the file is never written (it was being scanned, but not
150+
actively edited, so the file contents wasn't written back).
151+
152+
The autogenerated id must be a valid
153+
[CSS identifier](https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/es/ident).
154+
Autogenerated ids must be checked to make sure they don't collide with an
155+
existing id.
156+
157+
Design
158+
------
159+
160+
The cache is a single plain-data structure, shared as an `Arc<Mutex<Cache>>`;
161+
all consistency comes from that one lock, so no per-item locking (and therefore
162+
no lock ordering) is needed. Items refer to each other by key -- files by path,
163+
targets and fragments by id -- rather than by `Arc`/`Weak` pointers. This keeps
164+
the structure acyclic, `Send`, and (in the future) serializable, and avoids
165+
garbage-collecting stale weak references.
166+
167+
Cached state always converges by design. Each layer depends only on
168+
lower-numbered layers:
169+
170+
1. `Target::inner_html` (including a gather element's) and `Target::gather_ids`
171+
are pure functions of the source (pre-hydration) -- they depend on nothing.
172+
2. `<xref>` hydration depends on layer 1 only.
173+
3. `Fragment::content` depends on the source plus layer 2.
174+
4. `<fragment>` backlink hydration depends on layer 1 only: it renders, for each
175+
gather element which lists this fragment, that element's containing file,
176+
`id`, and inner HTML. Critically, this output is *excluded* from
177+
`Fragment::content` (layer 3); were it included, a gather element and each
178+
fragment it lists would outdate one another forever.
179+
5. Gather-list hydration depends on layers 1 and 3.
180+
181+
Nothing reads layers 2, 4, or 5, so propagation terminates.
182+
183+
Updating the cache is a two-phase process:
184+
185+
1. Collect: while walking a file's DOM, record all cacheable facts
186+
(`FileFacts`) -- targets, cross-references, fragments, and gather elements --
187+
without touching the cache. This keeps the non-`Send` DOM types out of the
188+
cache and off its critical section. Fact collection ignores content expanded
189+
by the cache (the contents of `xref` and `fragment` tags; the content
190+
following a gather tag).
191+
2. Commit: `Cache::commit_file` applies the facts in one transaction, diffing
192+
them against the file's previous state to compute the set of files outdated
193+
by these changes.
194+
195+
Item 2 requires the cache to track dependencies of an item, so that files
196+
containing these dependencies can be Outdated. Dependencies are tracked at file
197+
granularity, on the id rather than on the item defining it:
198+
`IdEntry::dependents` is the set of files whose rendered output depends on this
199+
id, by any means -- an `<xref ref="id">` or a `data-gather` list naming it. One
200+
uniform set suffices in all three `IdState`s because the typical action taken on
201+
a dependent is marking its containing file outdated (fragments also used their
202+
dependents to generate "See x" links).
203+
204+
Item 2 also requires the cache to define what constitutes a difference which
205+
would trigger outdating dependencies. The state which is checked for differences
206+
is:
207+
208+
* Target/gather element: type (a target/gather element), path of the containing
209+
file, id, IdState, inner HTML, and gather\_ids.
210+
* Fragment: type (a fragment), path of the containing file, id, IdState, and
211+
contents.
212+
213+
This also makes indirection (gather A includes fragment B, whose content
214+
cross-references target C) work without extra machinery: a change to C outdates
215+
B's file; reprocessing B's file changes B's content, which outdates A's file.
216+
217+
### Gather elements: the reverse edge
218+
219+
`dependents` alone is not enough for gather elements, because the reference runs
220+
in both directions: a gather element renders the *contents* of each fragment it
221+
lists, and each of those fragments renders a *backlink* to the gather element
222+
(see layer 4 above). The set of gather elements referencing a fragment is
223+
therefore not merely bookkeeping -- it is observable output in the fragment's
224+
own file.
225+
226+
This is why the outgoing references of gather elements cannot be maintained by
227+
the unlink-all/relink-all round trip used for cross-references (see
228+
`commit_file`): that round trip destroys the very information a diff would need.
229+
`commit_file` instead diffs `Target::gather_ids` explicitly, and applies this
230+
rule:
231+
232+
> When a gather element `G` is committed, outdate the file defining each id in
233+
> the symmetric difference of `G`'s old and new `gather_ids`; if `G`'s inner
234+
> HTML also changed, outdate the file defining each id in the union instead,
235+
> since the backlink text those fragments render comes from it.
236+
237+
A gather element moving to a different file needs no special case: the old
238+
file's commit sees the element deleted (new `gather_ids` empty) and the new
239+
file's commit sees it added, so both ends of the move outdate the fragments.
240+
241+
Propagation terminates here for the reason given under layer 4: rebuilding a
242+
fragment's file updates its backlinks, but backlinks are excluded from
243+
`Fragment::content`, so nothing further is outdated.
244+
245+
An id referenced before (or without) being defined has an `IdEntry` whose state
246+
is `IdState::Missing`; `dependents` holds the files waiting on it. When the id
247+
later appears, the state transition marks those files outdated, and they simply
248+
remain dependents; when a defined id disappears, the state returns to `Missing`
249+
and the dependents are again waiters. An `IdEntry` which is `Missing` with no
250+
dependents carries no information and is removed.
251+
252+
Duplicate ids are never renamed (file timestamps can't reliably identify the
253+
original, and renaming would silently modify user content). Instead, duplicates
254+
are reported as errors to the user.
255+
256+
### Misc
257+
258+
All file names (stored as `PathBuf`) must be canonicalized absolute paths.
259+
260+
HTML stored as `Target` inner HTML or in `Fragment` contents requires cleaning:
261+
262+
* To avoid duplicate IDs, all `id` attributes should be stripped.
263+
* Note that images, URLs, etc. may not work if the referring path in their new
264+
location isn't valid; these simply aren't supported.
265+
* Inner HTML must allow only the
266+
[permitted content for an `<a>` element](https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/Elements/a#technical_summary).
267+
268+
Watcher and walker
269+
------------------
270+
271+
Need a watcher/walker (WW). Each time a project is found, the code must ask for
272+
the cache for this project path. The WW maintains a connection id -> project
273+
path mapping. If the current request doesn't change the mapping, simply return
274+
the existing cache. If the mapping changes, then update the WW thingy. The map
275+
must also be updated when a connection is closed (to remove an existing
276+
mapping).
277+
278+
The WW thingy is another map from project path -> (cache, vec of project paths
279+
that are contained with this path, including the same path \[when multiple
280+
connections edit the same project\]). Operations:
281+
282+
* Insert: to insert a new project path, walk all existing top-level project
283+
paths. If the new project path is contained within any of these, add it to the
284+
appropriate list. Otherwise, add a new entry.
285+
* Delete: find the path my looking fir

0 commit comments

Comments
 (0)