Skip to content

feat: Support inline image(BI) - #89

Open
kvii wants to merge 1 commit into
ledongthuc:masterfrom
kvii:pr_BI
Open

kvii wants to merge 1 commit into
ledongthuc:masterfrom
kvii:pr_BI

Conversation

@kvii

@kvii kvii commented Sep 17, 2026 •

Copy link
Copy Markdown

Close: #30

Saving image data as string because I don't want to break current API.
Add Value.RawBytes for getting image data conveniently.

An example of extracting inline image from 99.pdf:

bb, _ := os.ReadFile("99.pdf")
br := bytes.NewReader(bb)
pr, _ := pdf.NewReader(br, br.Size())

var i int
contents := pr.Page(1).V.Key("Contents")
pdf.Interpret(contents, func(stk *pdf.Stack, op string) {
	if op != "BI" {
		return
	}

	bi := stk.Pop()
	w := int(bi.Key("W").Int64()) // 95
	h := int(bi.Key("H").Int64()) // 96
	// decode Image Data
	id := bi.Key("ID").RawBytes()
	br := bytes.NewReader(id)
	rd := ccitt.NewReader(br, ccitt.MSB, ccitt.Group4, w, h, nil)
	bs, _ := io.ReadAll(rd)

	// write image
	img := image.NewGray(image.Rect(0, 0, w, h))
	for y := range h {
		for x := range w {
			idx := y*((w+7)/8) + x/8
			if idx >= len(bs) {
				break
			}
			bitIdx := 7 - (x % 8)
			bit := (bs[idx] >> bitIdx) & 1
			img.SetGray(x, y, color.Gray{Y: bit * 255}) // 0->0, 1->255
		}
	}
	var buf bytes.Buffer
	err = png.Encode(&buf, img)
	if err != nil {
		t.Fatal(err)
	}
	out := buf.Bytes()

	i++
	os.WriteFile(fmt.Sprintf("out%d.png", i), out, 0666)
})

Cwooper added a commit to Cwooper/pdf that referenced this pull request Oct 5, 2026
The binary data after an inline image's ID was read as tokens, so an
unbalanced "(" in it opened a string that swallowed the rest of the
page, and other bytes counted as malformed operands. Interpret now skips
the data up to the first EI between whitespace, and its callers still
see the ID and EI operators. This is only a skip: it leaves room for
parsing the image dictionary and data, as upstream PR ledongthuc#89 proposes.
Cwooper added a commit to Cwooper/pdf that referenced this pull request Oct 5, 2026
The binary data after an inline image's ID was read as tokens, so an
unbalanced "(" in it opened a string that swallowed the rest of the
page, and other bytes counted as malformed operands. Interpret now skips
the data up to the first EI between whitespace, and its callers still
see the ID and EI operators. This is only a skip: it leaves room for
parsing the image dictionary and data, as upstream PR ledongthuc#89 proposes.
Cwooper added a commit to Cwooper/pdf that referenced this pull request Oct 6, 2026
The binary data after an inline image's ID was read as tokens, so an
unbalanced "(" in it opened a string that swallowed the rest of the
page, and other bytes counted as malformed operands. Interpret now skips
the data up to the first EI between whitespace, and its callers still
see the ID and EI operators. This is only a skip: it leaves room for
parsing the image dictionary and data, as upstream PR ledongthuc#89 proposes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

crash when encountering some CJK text amongst English

1 participant