History

jedarden 39ca6a3552 feat(pdftract-2b7ff): implement image_coverage_fraction signal evaluator Add image_coverage_fraction signal evaluator that computes the union image coverage fraction from individual image XObject areas. - Computes total image coverage as sum of image_xobject_areas - Divides by page area (width * height) to get coverage fraction - Clamps to [0.0, 1.0] to handle overlapping images (defensive) - Returns Some(Vote::scanned(0.85)) if fraction > 0.85 Implementation uses sum for simplicity (overestimates coverage when images overlap), which is acceptable for the 0.85 threshold as it's a conservative signal. Can be revisited with Klee's algorithm for greater accuracy if needed. Acceptance criteria PASS: ✓ Page with one image covering 90% area → Some(Vote { 0.85, Scanned }) ✓ Page with multiple small images totaling 50% → None (below threshold) ✓ Page with no images → None ✓ Coverage clamped to 1.0 on overlapping images Also includes pre-existing infrastructure: - tr3_op_count field in PageContext - image_xobject_areas field in PageContext - all_tr3_with_full_page_image function - CharDensityRatioSignal evaluator These were necessary dependencies for the new evaluator to function. Refs: Plan section Phase 5.1.2, coordinator pdftract-22p		2026-05-31 23:42:38 -04:00
..
benches	feat(pdftract-1z0qt): add encryption verification note	2026-05-28 08:09:53 -04:00
bin	fix(pyo3): correct extract_text_fn call in extract_markdown stub	2026-05-28 20:28:25 -04:00
build	feat(pdftract-1xf4d): implement TH-06 supply-chain gate	2026-05-26 17:31:13 -04:00
examples	wip: AcroForm improvements, debug tooling, test corpus, and fixture updates	2026-05-30 09:48:14 -04:00
proptest-regressions/parser/lexer	feat(pdftract-1jjn): implement PDF numeric literal lexer with full edge case support	2026-05-23 23:17:04 -04:00
scripts	wip: intermediate state from previous work	2026-05-29 08:25:23 -04:00
src	feat(pdftract-2b7ff): implement image_coverage_fraction signal evaluator	2026-05-31 23:42:38 -04:00
tests	fix(test): add error handling for missing fixture paths	2026-05-31 14:12:44 -04:00
__test__.pdf	feat(pdftract-15pz8): implement multi-process safe cache operations	2026-05-23 05:31:11 -04:00
build.rs	fix(pdftract-4pnmd): build.rs doc comment format string parsing	2026-05-28 14:36:45 -04:00
Cargo.toml	feat(profiles): add profile infrastructure and initial fixtures	2026-05-31 15:10:51 -04:00
pdftract-core.cdx.json	feat(pdftract-67tm8): implement MCP stdio transport with integration tests	2026-05-23 00:16:42 -04:00
README.md	docs(pdftract-49f8): establish Cargo.lock policy and documentation	2026-05-20 18:13:14 -04:00

README.md

pdftract-core

The core Rust library for PDF text extraction. This crate provides the parsing, layout analysis, font encoding recovery, and text extraction primitives used by the CLI (pdftract-cli) and Python bindings (pdftract-py).

Cargo.lock Policy

This workspace checks in Cargo.lock at the repository root. This is unconventional for library crates—the Cargo Book historically suggested that only binary crates should check in lockfiles, allowing library consumers to resolve their own dependency versions.

pdftract departs from this convention for release reproducibility:

SLSA Level 3 provenance requires that every milestone tag produces byte-identical artifacts across builds. Without a checked-in lockfile, two runs of cargo build on the same commit can resolve different transitive dependency versions, producing different binary hashes.
Multi-output artifacts—this workspace produces Rust crates (pdftract-core, pdftract-cli), Python wheels (pdftract-py), and Docker images. All must be built from the same dependency tree.
Supply-chain security—the lockfile pins checksums for all transitive dependencies, enabling cargo audit to detect yanked or compromised crates.
Downstream consumers can still ignore the lockfile if needed. Cargo allows cargo build --frozen with a local lockfile override, or consumers can vendor the crate with their own dependency resolution.

The tradeoff—occasional merge conflicts when PRs update overlapping dependencies—is worth the guarantee of reproducible releases. See CONTRIBUTING.md for the lockfile-update workflow.

Modules

parser: PDF spec parsing (xref, trailer, object streams, indirect references)
font: Font encoding recovery, glyph name lookup, fingerprinting
layout: Page layout analysis, region segmentation, reading order
extract: Text extraction with provenance (bounding boxes, confidence scores)
ocr: Tesseract integration for raster pages

Usage

use pdftract_core::{extract_text, ExtractOptions};

let options = ExtractOptions::default();
let result = extract_text("document.pdf", &options)?;
println!("{}", result.text);