Skip to content

fix(docreader): make DOCX header extraction optional - #3914

Open
helloxjade wants to merge 1 commit into
Tencent:mainfrom
helloxjade:fix/docx-optional-headers
Open

helloxjade wants to merge 1 commit into
Tencent:mainfrom
helloxjade:fix/docx-optional-headers

Conversation

@helloxjade

Copy link
Copy Markdown

Description

DOCX parsing currently returns only body content, so text that appears in a document header cannot be retrieved. Add an opt-in docx_include_headers parser rule for the builtin and MarkItDown engines, exposed in knowledge-base parser settings and the upload confirmation dialog. The default remains disabled.

When enabled, parsing prepends header text and tables, includes active first-page and even-page headers, and extracts inherited header parts only once. The python-docx fallback also honors the option. The option covers text and tables; header images and external AnyDoc processing are outside its supported extraction behavior.

Type of Change

  • 🐛 Bug fix
  • 🧪 Test
  • 📚 Documentation update

Related Issue

Fixes #3849

Testing

  • Reproduced missing header text in the builtin and MarkItDown paths with an in-memory DOCX before the fix.
  • python -m unittest docreader.tests.test_docx_headers docreader.tests.test_docx_tables docreader.tests.test_docx_merge -q — 25 tests passed.
  • go test ./internal/application/service -run 'TestApplyParserRuleOverrides|TestResolveProcessConfig' -count=1 — passed.
  • npm test -- src/views/knowledge/settings/KBParserSettings.test.ts src/views/knowledge/components/UploadConfirmDialog.test.ts src/i18n/localeKeyAudit.test.ts — 22 tests passed.
  • npm run type-check and npm run build — passed.
  • Browser check of the actual parser-settings component with fixture engine resources: checkbox initially unchecked; clicking it emits the DOCX rule with docx_include_headers: true.
  • git diff --check upstream/main...HEAD — passed (upstream is Tencent/WeKnora).
  • golangci-lint run --new-from-rev=upstream/main ./... could not run because golangci-lint is not installed. The full-repository maintainer gate was not run; validation was scoped to the changed components.

Checklist

  • Diff whitespace check passes against upstream main
  • Changed source files are formatted
  • Targeted tests for the changed packages/components pass
  • Diff-scoped lint passes where applicable (tool unavailable; see above)
  • Full-repository check limitations documented above
  • Self-reviewed the code
  • Added/updated tests covering the change
  • Updated related documentation

Screenshots / Recordings

Actual KBParserSettings component, rendered with fixture parser engines:

DOCX header extraction option

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: docx 使用anydoc markitdown docreader都无法解析docx的页眉

1 participant