Repository navigation
feat(find): harden search failure modes and blend ranking by popularity - #2457
Open
lanyouxize wants to merge 1 commit into
Open
lanyouxize wants to merge 1 commit into
lanyouxize wants to merge 1 commit into
Conversation
`skills find` had four defects that made the command unreliable and, in one case, actively misleading. 1. Failures were indistinguishable from "no matches". `searchSkillsAPI` returned `[]` for a non-2xx response, a JSON parse error and a network exception alike. The CLI then printed `No skills found for "<query>"`, so a 503 or a DNS failure was reported to the user as a factual search result. Failures now return a typed `SearchOutcome` (`ok | empty | timeout | server-error | network-error`); only `empty` legitimately means "no match". Non-`ok` outcomes print an actionable message on stderr, state explicitly that this is NOT proof of no match, and exit non-zero so callers can retry. 2. The search request had no timeout. `fetch(url)` could hang indefinitely. Sibling code (`download-source.ts`) already used `AbortSignal.timeout(30_000)`; the search path now uses an 8s bound (override via `SKILLS_SEARCH_TIMEOUT_MS`), which is deliberately tighter because searching is interactive. 3. Out-of-order responses could overwrite fresher results. Debouncing throttles the request *rate* but does not order the responses; a slow reply for an older query could land after and clobber a newer one. Each request now carries a monotonic sequence number and stale replies are discarded. Repeated queries also reuse a 60s TTL cache, so backspacing and retyping no longer re-hit the network. 4. Ranking discarded server relevance. The response is already relevance ordered, but the client re-sorted by raw `installs`, which promoted unrelated high-popularity skills. On a live `q=paper` capture this left 3/10 top results with "paper" in the name. Ranking is now a weighted blend (0.7 popularity / 0.3 relevance by default). Popularity is log-compressed because install counts span five orders of magnitude (15 - 452,294) and would otherwise swamp the relevance term entirely, collapsing the blend back into the old pure-installs sort. On the same capture the top-10 relevant density improves from 3 to 5. Success-path behaviour is unchanged: same endpoint, parameters and field names. Adds no dependencies. Tests: 29 new cases covering the outcome taxonomy (each failure mode asserted distinct from empty), sequence-guard ordering, cache TTL/eviction/owner isolation, and ranking calibration. The ranking suite is built on a real API capture and sweeps the weight grid so the 0.7/0.3 default is backed by data rather than intuition.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
skills findhad four defects. One of them was actively misleading: a failedsearch was reported to the user as "No skills found", i.e. a network outage was
presented as a factual search result.
[], same as "no matches"fetch(url)had no timeoutinstallsChanges
1. Typed failure taxonomy.
searchSkillsAPInow returnsSearchOutcome=ok | empty | timeout | server-error | network-error.Only
emptymeans "no match". Everything else prints an actionable message onstderr and exits non-zero.
Before:
After:
2. Request timeout. Added
AbortSignal.timeout(8_000)on the search path(override with
SKILLS_SEARCH_TIMEOUT_MS). The sibling download path alreadyused a 30s bound; 8s is deliberately tighter because searching is interactive and
the user has already concluded the tool is broken by then.
3. Response ordering + cache. Each request carries a monotonic sequence
number; replies whose sequence is stale are discarded. Debouncing only throttles
the request rate, it does not order responses. Repeated queries within 60s
reuse a TTL cache, so backspacing/retpying no longer re-hits the network.
4. Ranking blend. The API response is already relevance-ordered; the client
was discarding that by re-sorting on raw
installs. Ranking is now0.7 * popularity + 0.3 * relevance.Popularity is log-compressed (
log10(1+n)/log10(1e6)). This matters: installcounts span five orders of magnitude (15 to 452,294), so a raw-weight blend
would be fully swamped by the popularity term and collapse back into the old
pure-installs sort. The log makes the two terms comparable.
Measured on a live
q=papercapture:Weight choice is data-backed, not arbitrary
The ranking suite sweeps the weight grid:
Density saturates at 0.5-0.7; 0.7 is the point where density is maxed while
unrelated promotions stay bounded. 1.0 reproduces the old defect.
Compatibility
AbortSignal.timeoutrequires 17.3+; CI matrix is 22.20.0 / 24 / 26.SKILLS_API_URL/SKILLS_SEARCH_TIMEOUT_MSare opt-in overrides.Testing
29 new cases across two files:
search-hardening.test.ts(17) - each failure mode asserted distinctfrom
empty; AbortSignal always passed; never throws on network errors;sequence-guard ordering; cache TTL / eviction / owner isolation / case-insensitivity.
find-ranking.test.ts(12) - score bounds, monotonicity, determinism,input immutability, empty & single-element lists, weight-grid calibration on a
real API capture.
The ranking tests use a real API capture as fixture rather than mocks, so
they encode the actual server response shape.
Test suite was adversarially verified
A green suite proves nothing unless it can go red. Three injected defects, each
caught:
expected 3 to be greater than 3isCurrent()-> always true (guard disabled)expected true to be falseAll defects reverted;
grep -r "INJECTED BUG" src/returns nothing.Zero-regression evidence
Failure count is identical, and the delta is exactly the new tests.
Verified from a clean clone of the branch:
pnpm install --frozen-lockfile->tsc --noEmitPASS -> 34/34 tests PASS -> liveskills find paperreturnsranked results.