Repository navigation
Conversation
EPUBParser took the chapter order from book.get_table_of_contents(), which ebooklib's EpubBook does not have. The AttributeError was caught, so every book fell through to get_items(), which follows the manifest. The manifest may list files in any order; the spine is the reading order. Standard Ebooks' Pride and Prejudice, for example, lists its title page and imprint after chapter 61 in the manifest. Read the documents in spine order, followed by any the spine leaves out in manifest order, so the set of documents read is unchanged.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
EPUBParser._extract_contentmeant to order chapters by the table of contents, but that branch never runs:book.get_table_of_contents()does not exist on ebooklib'sEpubBook(checked on ebooklib 0.20, the version indocreader/uv.lock), so theAttributeErroris caught andtocis always[].epub.Linkobjects, which have noget_name(), so the loop would match nothing either way.Every book therefore falls back to
book.get_items(), which is manifest order. The manifest may list files in any order. The spine is the reading order. Real books differ:src/epub/content.opf)moby-dickaccessible_epub_3The fix reads the documents in spine order and then appends any document the spine leaves out, in manifest order. The set of documents read stays the same. Only their order changes, and the dead TOC code goes. The ZIP fallback for EPUBs that ebooklib cannot open is untouched.
This affects the builtin DocReader engine, which handles EPUB when it is selected for a knowledge base or when the server is built without anydoc (
preferAnydocWhenAvailableroutes EPUB to anydoc when it is linked). #3924 also edits this file, but only the image-alias helpers further down, so the two changes do not overlap.Type of Change
Related Issue
None filed.
Testing
Environment:
uv sync --project docreader --locked --no-dev --python 3.10.18(as indocreader.yml), macOS.New tests in
docreader/tests/test_epub_parser.pybuild EPUBs with ebooklib whose manifest order differs from the spine:test_chapters_follow_the_spine_not_the_manifest: manifestthree, title, one, two, spinetitle, one, two, three, plus an SVG cover in the spine that must stay out.test_documents_outside_the_spine_follow_it: a document missing from the spine is still read, after the spine.test_a_document_listed_twice_in_the_spine_is_read_once.With
main's parser the first two fail (AssertionError: Lists differ: [44, 72, 98, 14] != [14, 44, 72, 98]and[74, 48, 17] != [17, 48, 74]). Five mutants of the new function each fail at least one test: no spine loop, spine reversed, documents outside the spine dropped, the document filter removed (the SVG gets in), and no de-duplication.The one error is
test_ssrf_proxy.TestSSRFProxy.test_webkit_redirects_and_subresources_use_proxy, which needs the Playwright WebKit browser that CI installs and this machine does not have. It is unrelated to this change.Checklist
git diff --check origin/main...HEADpassesdocreader; the new code follows the file's existing style)golangci-lint run --new-from-rev=origin/main ./...) (no Go changes)website-docs/, Swagger annotations, etc.)