Repository navigation
Conversation
6 of 17 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
The builtin DocReader can successfully parse an EPUB while associating a chapter with the wrong embedded image. When
first/pic.png,second/pic.pngand a root-levelpic.pngcoexist,<img src="pic.png">infirst/chapter.xhtmlcan resolve to another resource because a basename alias is checked before the chapter-relative path.Resolve paths relative to the chapter first, retain basename matching as the final fallback, and prevent basename aliases from overwriting real archive-root paths. Both ebooklib parsing and the ZIP fallback use this logic. Regression tests generate EPUBs with distinct real PNGs and compare the bytes behind each Markdown reference.
The follow-up also preserves literal
#,?and percent-encoded-looking characters in archive paths. Archive entry names from EbookLib and ZIP are already decoded paths; treating them as URI references could truncate those names or decode them twice and select a different image. Only the HTML imagesrcis treated as a URI reference: remove its query and fragment before decoding once, then resolve it relative to the chapter. Additional byte-level regression cases exercise both EbookLib and the ZIP fallback.Type of Change
Related Issue
No separate issue filed. All-state EPUB and image/basename searches were refreshed on 2026-10-02; no directly matching existing issue or PR was found.
Testing
Latest PR head
On 2026-10-02, validated the current head
5f952f85843cf231068428e2ae66a484771bf4dewith the EPUB unit suite on Windows using CPython 3.10.18 and the existing installed dependencies. The production parser and test files were fetched at that exact commit, verified against their Git blob hashes, and executed from an isolated temporary source copy. Module paths and the parser registry were checked to ensure they used that same updated parser.subTestexecutions; 0 failures, 0 errors and 0 skips. The two newly added methods each cover 7 literal-path cases: 14 new subtests across EbookLib parsing and ZIP fallback. The remaining 5 subtests cover the original chapter-relative and image-order cases.git diff --check origin/main...HEAD: passes on this latest head, with a clean source working tree.The full Python suite, production gRPC harness, Go clients, full repository Go tests, formatting and lint were not rerun on this latest head. Their results below apply to the initial commit only.
Initial commit: historical Linux validation
Validated the initial PR commit
4d662b9c28e82d0f71a0f75e1405b82b3be22898on GitHub-hosted Ubuntu Linux using CPython 3.10.18, lockeddocreader/uv.lockdependencies and Go 1.26.0. The validation workflow is on a separate fork branch and is not part of this PR.python -m compileall -q docreader: passes.python -m unittest discover -s docreader/tests -p 'test_*.py' -v: 266 tests, 254 pass, 12 skip, no failures or errors. Eleven skips require unavailable fixtures; one requires ImageMagick. All seven EPUB test methods pass.bccb4b151bae403508da77fbb174efc79dc47c1a, with the test registry pointing to that same parser: six expected failures across the new regression cases; all four original test methods pass.DocReaderServicerverifies the actual PNG bytes for unaryRead, streamingReadStream, andReadwith ZIP fallback. All three paths pass; no model calls.go test ./docreader/client ./docreader/proto -count=1 -v, with a live DocReader service: TestReadURL and TestReadFile pass, no skips;protohas no test files. The URL test readshttps://example.com.make test(go test -v ./...): 115 packages pass, no failed packages or tests; 46 tests/subtests skip under existing conditions for PostgreSQL, model credentials, native browser/sandbox tools, SQLite FTS5 and other optional integration resources. Skipped cases are not counted as passed.make fmt: passes with no Go diff.git diff --checkagainst the PR baseline passes.make lint, golangci-lint v2.12.2: does not pass; reports 246 findings in existing Go source. This PR changes no Go source, lint configuration or dependencies. Lint findings remain explicitly reported and were not fixed or suppressed as part of this Python-only change.Inspect these historical logs for the initial commit's underlying test outcomes, rather than relying only on the workflow's aggregate green status:
These historical Linux checks were run by the contributor for the initial commit. Tencent's upstream DocReader workflow still requires maintainer authorization. Frontend checks and a full deployed upload/index/query end-to-end run are outside the validation performed for this parser fix. No unrelated source or test assertions were changed.
Checklist
git diff --check origin/main...HEADpassesAI assistance
Codex assisted with investigation, reproduction, implementation, tests and PR preparation.