|
| 1 | +Cache spec |
| 2 | +========== |
| 3 | + |
| 4 | +The cache stores the location (file name and ID) and contents (inner HTML or |
| 5 | +code/doc blocks for fragments) of a target. Targets are HTML elements (excluding |
| 6 | +`<fragment>`s) with an ID. |
| 7 | + |
| 8 | +The goal of the cache is to support cross-references and gather elements, and to |
| 9 | +ensure that all IDs are unique within a project. This means that |
| 10 | +cross-references and gather elements persist across moving or renaming files, |
| 11 | +since the IDs will be found in the cache. |
| 12 | + |
| 13 | +Non-project files support a subset of this functionality: the "project" consists |
| 14 | +only of the current file. Only targets, gather elements, cross-references, and |
| 15 | +fragments to items within the file work as expected; references to other files |
| 16 | +do not. |
| 17 | + |
| 18 | +The cache reflects data read directly from disk/IDE; content edited in the |
| 19 | +Client does not update the cache until it's written to disk/IDE, at which point |
| 20 | +the cache much re-process this unknown file. Since the Client is designed around |
| 21 | +an autosave principle which updates disk/IDE regularly, there's little gap |
| 22 | +between the two. |
| 23 | + |
| 24 | +The cache hydrates data to the Client; the dehydration routes are responsible |
| 25 | +for removing all cache hydration artifacts. |
| 26 | + |
| 27 | +Cross references |
| 28 | +---------------- |
| 29 | + |
| 30 | +A `<xref ref="id"></xref>` is a cross reference to a `Target` or a gather |
| 31 | +element. The `id` specifies the destination; the cache then hydrates the |
| 32 | +contents based on the location and contents of the target of the provided `id` |
| 33 | +to e.g. `<xref ref="id" contenteditable="false"><a |
| 34 | +href="../path/to/page#id">Inner HTML from target</a></xref>`; this hydrated form |
| 35 | +is only present in the Client. |
| 36 | + |
| 37 | +Details: |
| 38 | + |
| 39 | +* This element does not allow an `id` attribute. |
| 40 | + |
| 41 | +* The inner HTML is always taken before cache hydration, to prevent circular |
| 42 | + dependencies: |
| 43 | + |
| 44 | + ```html |
| 45 | + <h1 id="a">See <xref ref="b"></xref></h1> |
| 46 | + <!-- file a --> |
| 47 | + <h1 id="b">See <xref ref="a"></xref></h1> |
| 48 | + <!-- file b --> |
| 49 | + ``` |
| 50 | + |
| 51 | + Including cache hydration would cause the inner HTML to be updated each time |
| 52 | + the file is processed, outdating the other file. |
| 53 | + |
| 54 | +* If the `id` referred to isn't found or refers to a duplicate id, the inner |
| 55 | + text is instead an appropriate error message. |
| 56 | + |
| 57 | +* If the cross reference is to a gather element, the text is the gather |
| 58 | + element's inner HTML, not the gathered code/doc blocks. |
| 59 | + |
| 60 | +Fragments and gather elements |
| 61 | +----------------------------- |
| 62 | + |
| 63 | +A gather element such as `<h3 id="bar" data-gather="id1 id2...">Bazzy |
| 64 | +things</h3>` is a `Target` with the `data-gather` attribute. It becomes a list |
| 65 | +of the contents of fragments it refers to after hydration by the cache. An |
| 66 | +example fragment tag, after cache hydration: `<fragment id="id1" |
| 67 | +contenteditable="false">See <a href="path/to/gather#bar">Bazzy things</a>, <a |
| 68 | +href="path/to/another/gather#zap">Zappy things</a></fragment>`. A fragment's |
| 69 | +content by default includes the contents of the current doc block and the |
| 70 | +contents of the following code/doc block; fragments are not allowed in Markdown |
| 71 | +documents (in the case, the fragment contents consist of an error message). |
| 72 | +Fragments may include the `following` attribute to enclose a specific number of |
| 73 | +the following code/doc blocks; for example, `<fragment id="bar" |
| 74 | +following="3"></fragment>` includes the current doc block along with the next 3 |
| 75 | +code/doc blocks; `following` must be a whole number. |
| 76 | + |
| 77 | +Details: |
| 78 | + |
| 79 | +* Fragment contents may not include a gather element; in this case, the gather |
| 80 | + element list of contents will simply include an error message. |
| 81 | +* Fragments do support indirection: gather element A includes contents from |
| 82 | + fragment B, which contains a cross reference to target C. Changes to target C |
| 83 | + makes B and A outdated. |
| 84 | +* If a gather element refers to an `id` that is a `Target` or a `GatherElement`, |
| 85 | + not a `Fragment`, the resulting output for this in the list of fragments is an |
| 86 | + error message. |
| 87 | +* If a gather element refers to an `id` that wasn't found or is a duplicate, its |
| 88 | + contents will be replaced by an error message. |
| 89 | +* Fragments store an HTML rendering of the code and doc blocks they contain, |
| 90 | + excluding the content produced by hydrating the `<fragment>` tags, to avoid |
| 91 | + duplication and circular dependencies. See layer 4 under `Design`: this |
| 92 | + exclusion is what keeps a gather element and the fragments it lists from |
| 93 | + outdating each other forever. The HTML rendering of a fragment reproduces the |
| 94 | + layout of the source it came from: each doc block includes its indent, each |
| 95 | + line of a code block is preceded by that line's number, and the two are |
| 96 | + aligned -- a doc block and a line of code indented equally in the source begin |
| 97 | + in the same column, with the line numbers in a gutter of their own to the left |
| 98 | + of both. |
| 99 | +* The backlinks a `<fragment>` hydrates to are derived from the gather elements |
| 100 | + which list it: for each such element, its containing file, its `id`, and its |
| 101 | + inner HTML (the link text). A fragment's rendered output therefore depends on |
| 102 | + the gather elements which reference it -- the reverse of the direction in |
| 103 | + which references are written; see `Design`. |
| 104 | +* A `<fragment following=0>` is valid; it contains only the current doc block. |
| 105 | +* The `following` attribute is clamped if it would exceed the number of code/doc |
| 106 | + blocks in the document. |
| 107 | +* If the `following` value cannot be parsed to a whole number, an error message |
| 108 | + replaces the fragment content. |
| 109 | +* A gather element, as a type of `Target`, requires an `id`. A `data-gather` |
| 110 | + attribute on an element without an `id` produces an error message which |
| 111 | + requests the missing `id`. |
| 112 | +* All ids must be valid |
| 113 | + [CSS identifiers](https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/es/ident) |
| 114 | + per |
| 115 | + [MDN recommendations](https://developer.mozilla.org/en-US/docs/Web/HTML/rence/Global_attributes/id). |
| 116 | + Invalid ids produce error messages in the hydrated tag content. |
| 117 | + |
| 118 | +Example hydration of the gather tag `<h3 id="bar" data-gather="id1 id2...">Bazzy |
| 119 | +things</h3>`: |
| 120 | + |
| 121 | +```html |
| 122 | +<h3 class="cc-gather" id="bar" data-gather="id1 id2...">Bazzy things</h3> |
| 123 | +<div class="cc-gather-items" contenteditable="false"> |
| 124 | + <p class="cc-gather-item-link"> |
| 125 | + From <a href="link/to/first/tag#id1">Path to file</a>: |
| 126 | + </p> |
| 127 | + (first item content) ... |
| 128 | + <p class="cc-gather-item-link"> |
| 129 | + From <a href="link/to/last/tag#idn">Path to file</a>: |
| 130 | + </p> |
| 131 | + (last item content) |
| 132 | +</div> |
| 133 | +``` |
| 134 | + |
| 135 | +Search |
| 136 | +------ |
| 137 | + |
| 138 | +The cache supports searching the (cleaned) inner HTML of all `Target`s and |
| 139 | +gather elements; search does not include `Fragment` contents. |
| 140 | + |
| 141 | +### Auto-assignment of ids |
| 142 | + |
| 143 | +If `id="*"` on either a fragment, target, or gather element, the cache replaces |
| 144 | +this with an random autogenerated `id` placed in the resulting HTML contents, |
| 145 | +but this new `id` is not yet recorded in the cache. The file is then marked as |
| 146 | +`Unknown` when it is saved; when re-read, this `id` is then incorporated into |
| 147 | +the cache. This helps avoid cases where the cache and file contents become |
| 148 | +unsynchronized: if the `id` is placed in the cache before the write and the |
| 149 | +write fails, or if the file is never written (it was being scanned, but not |
| 150 | +actively edited, so the file contents wasn't written back). |
| 151 | + |
| 152 | +The autogenerated id must be a valid |
| 153 | +[CSS identifier](https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/es/ident). |
| 154 | +Autogenerated ids must be checked to make sure they don't collide with an |
| 155 | +existing id. |
| 156 | + |
| 157 | +Design |
| 158 | +------ |
| 159 | + |
| 160 | +The cache is a single plain-data structure, shared as an `Arc<Mutex<Cache>>`; |
| 161 | +all consistency comes from that one lock, so no per-item locking (and therefore |
| 162 | +no lock ordering) is needed. Items refer to each other by key -- files by path, |
| 163 | +targets and fragments by id -- rather than by `Arc`/`Weak` pointers. This keeps |
| 164 | +the structure acyclic, `Send`, and (in the future) serializable, and avoids |
| 165 | +garbage-collecting stale weak references. |
| 166 | + |
| 167 | +Cached state always converges by design. Each layer depends only on |
| 168 | +lower-numbered layers: |
| 169 | + |
| 170 | +1. `Target::inner_html` (including a gather element's) and `Target::gather_ids` |
| 171 | + are pure functions of the source (pre-hydration) -- they depend on nothing. |
| 172 | +2. `<xref>` hydration depends on layer 1 only. |
| 173 | +3. `Fragment::content` depends on the source plus layer 2. |
| 174 | +4. `<fragment>` backlink hydration depends on layer 1 only: it renders, for each |
| 175 | + gather element which lists this fragment, that element's containing file, |
| 176 | + `id`, and inner HTML. Critically, this output is *excluded* from |
| 177 | + `Fragment::content` (layer 3); were it included, a gather element and each |
| 178 | + fragment it lists would outdate one another forever. |
| 179 | +5. Gather-list hydration depends on layers 1 and 3. |
| 180 | + |
| 181 | +Nothing reads layers 2, 4, or 5, so propagation terminates. |
| 182 | + |
| 183 | +Updating the cache is a two-phase process: |
| 184 | + |
| 185 | +1. Collect: while walking a file's DOM, record all cacheable facts |
| 186 | + (`FileFacts`) -- targets, cross-references, fragments, and gather elements -- |
| 187 | + without touching the cache. This keeps the non-`Send` DOM types out of the |
| 188 | + cache and off its critical section. Fact collection ignores content expanded |
| 189 | + by the cache (the contents of `xref` and `fragment` tags; the content |
| 190 | + following a gather tag). |
| 191 | +2. Commit: `Cache::commit_file` applies the facts in one transaction, diffing |
| 192 | + them against the file's previous state to compute the set of files outdated |
| 193 | + by these changes. |
| 194 | + |
| 195 | +Item 2 requires the cache to track dependencies of an item, so that files |
| 196 | +containing these dependencies can be Outdated. Dependencies are tracked at file |
| 197 | +granularity, on the id rather than on the item defining it: |
| 198 | +`IdEntry::dependents` is the set of files whose rendered output depends on this |
| 199 | +id, by any means -- an `<xref ref="id">` or a `data-gather` list naming it. One |
| 200 | +uniform set suffices in all three `IdState`s because the typical action taken on |
| 201 | +a dependent is marking its containing file outdated (fragments also used their |
| 202 | +dependents to generate "See x" links). |
| 203 | + |
| 204 | +Item 2 also requires the cache to define what constitutes a difference which |
| 205 | +would trigger outdating dependencies. The state which is checked for differences |
| 206 | +is: |
| 207 | + |
| 208 | +* Target/gather element: type (a target/gather element), path of the containing |
| 209 | + file, id, IdState, inner HTML, and gather\_ids. |
| 210 | +* Fragment: type (a fragment), path of the containing file, id, IdState, and |
| 211 | + contents. |
| 212 | + |
| 213 | +This also makes indirection (gather A includes fragment B, whose content |
| 214 | +cross-references target C) work without extra machinery: a change to C outdates |
| 215 | +B's file; reprocessing B's file changes B's content, which outdates A's file. |
| 216 | + |
| 217 | +### Gather elements: the reverse edge |
| 218 | + |
| 219 | +`dependents` alone is not enough for gather elements, because the reference runs |
| 220 | +in both directions: a gather element renders the *contents* of each fragment it |
| 221 | +lists, and each of those fragments renders a *backlink* to the gather element |
| 222 | +(see layer 4 above). The set of gather elements referencing a fragment is |
| 223 | +therefore not merely bookkeeping -- it is observable output in the fragment's |
| 224 | +own file. |
| 225 | + |
| 226 | +This is why the outgoing references of gather elements cannot be maintained by |
| 227 | +the unlink-all/relink-all round trip used for cross-references (see |
| 228 | +`commit_file`): that round trip destroys the very information a diff would need. |
| 229 | +`commit_file` instead diffs `Target::gather_ids` explicitly, and applies this |
| 230 | +rule: |
| 231 | + |
| 232 | +> When a gather element `G` is committed, outdate the file defining each id in |
| 233 | +> the symmetric difference of `G`'s old and new `gather_ids`; if `G`'s inner |
| 234 | +> HTML also changed, outdate the file defining each id in the union instead, |
| 235 | +> since the backlink text those fragments render comes from it. |
| 236 | +
|
| 237 | +A gather element moving to a different file needs no special case: the old |
| 238 | +file's commit sees the element deleted (new `gather_ids` empty) and the new |
| 239 | +file's commit sees it added, so both ends of the move outdate the fragments. |
| 240 | + |
| 241 | +Propagation terminates here for the reason given under layer 4: rebuilding a |
| 242 | +fragment's file updates its backlinks, but backlinks are excluded from |
| 243 | +`Fragment::content`, so nothing further is outdated. |
| 244 | + |
| 245 | +An id referenced before (or without) being defined has an `IdEntry` whose state |
| 246 | +is `IdState::Missing`; `dependents` holds the files waiting on it. When the id |
| 247 | +later appears, the state transition marks those files outdated, and they simply |
| 248 | +remain dependents; when a defined id disappears, the state returns to `Missing` |
| 249 | +and the dependents are again waiters. An `IdEntry` which is `Missing` with no |
| 250 | +dependents carries no information and is removed. |
| 251 | + |
| 252 | +Duplicate ids are never renamed (file timestamps can't reliably identify the |
| 253 | +original, and renaming would silently modify user content). Instead, duplicates |
| 254 | +are reported as errors to the user. |
| 255 | + |
| 256 | +### Misc |
| 257 | + |
| 258 | +All file names (stored as `PathBuf`) must be canonicalized absolute paths. |
| 259 | + |
| 260 | +HTML stored as `Target` inner HTML or in `Fragment` contents requires cleaning: |
| 261 | + |
| 262 | +* To avoid duplicate IDs, all `id` attributes should be stripped. |
| 263 | +* Note that images, URLs, etc. may not work if the referring path in their new |
| 264 | + location isn't valid; these simply aren't supported. |
| 265 | +* Inner HTML must allow only the |
| 266 | + [permitted content for an `<a>` element](https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/Elements/a#technical_summary). |
| 267 | + |
| 268 | +Watcher and walker |
| 269 | +------------------ |
| 270 | + |
| 271 | +Need a watcher/walker (WW). Each time a project is found, the code must ask for |
| 272 | +the cache for this project path. The WW maintains a connection id -> project |
| 273 | +path mapping. If the current request doesn't change the mapping, simply return |
| 274 | +the existing cache. If the mapping changes, then update the WW thingy. The map |
| 275 | +must also be updated when a connection is closed (to remove an existing |
| 276 | +mapping). |
| 277 | + |
| 278 | +The WW thingy is another map from project path -> (cache, vec of project paths |
| 279 | +that are contained with this path, including the same path \[when multiple |
| 280 | +connections edit the same project\]). Operations: |
| 281 | + |
| 282 | +* Insert: to insert a new project path, walk all existing top-level project |
| 283 | + paths. If the new project path is contained within any of these, add it to the |
| 284 | + appropriate list. Otherwise, add a new entry. |
| 285 | +* Delete: find the path my looking fir |
0 commit comments