Receipt Forgery Detection Near Random? Real-World Struggles and Solutions in Document Image Forensics

Thesis case study: receipt tampering detection AUC near random, exposing classic failure modes in small-sample document forensics.
A student tried EfficientNet, BERT, numerical consistency checks, multimodal fusion, YOLO, and synthetic data augmentation on the ICDAR 2023 receipt forensics dataset — yet ROC-AUC stayed between 0.46 and 0.54. The article diagnoses the root causes: tampering signals are highly local and sparse; global pooling dilutes them entirely; synthetic data introduces artificial artifacts causing overfitting; and DocTamper zero-shot transfer fails due to severe domain shift. Recommended paths forward include reframing the task as patch-level anomaly detection, using self-supervised pre-training to establish normal receipt priors, treating numerical logic consistency as a standalone strong signal, and running diagnostic ablation experiments to distinguish data insufficiency from method failure.
A computer vision student shared their thesis predicament on Reddit: working with the ICDAR 2023 Find It Again! dataset to detect forged receipts and localize tampered regions, they found that no matter how they adjusted the model, ROC-AUC stubbornly hovered between 0.46 and 0.54 — essentially random guessing. This case is remarkably representative, exposing nearly every classic pitfall in small-sample document forensics tasks.

Why Receipt Forgery Detection Is So Hard
Receipt tampering detection is fundamentally different from standard image classification. The original poster identified several core challenges: the dataset is small and class-imbalanced; many forgeries change only one or two digits; the tampered regions are extremely small; and models tend to perform reasonably on training data while failing to generalize to unseen receipts.
The most critical tension here is signal sparsity. When only a few pixels' worth of digits are altered on a receipt, the image's global features carry almost no discriminative information. Using EfficientNet-B0 for image-level binary classification is like asking a model to find a pinhole-sized anomaly in an entire page of text — the global pooling operation completely dilutes that faint signal. This explains why the "frozen visual features" approach yielded a ROC-AUC of only 0.482: the model essentially learned nothing.
The ICDAR 2023 Find It Again! dataset is a specialized forensics benchmark released by the International Conference on Document Analysis and Recognition (ICDAR), focused on receipt image authenticity verification and tampered region localization. It is a small-scale competition dataset in the document image forensics domain. Unlike large-scale general vision datasets such as ImageNet with millions of samples, such datasets typically contain only hundreds to thousands of images, and positive samples (tampered receipts) are often far outnumbered by negatives — posing serious challenges to the statistical foundations of supervised learning. ROC-AUC (Area Under the Receiver Operating Characteristic Curve) is the standard metric for evaluating a binary classifier's overall discriminative ability across different decision thresholds. It ranges from 0 to 1, where 0.5 corresponds to random guessing and 1.0 is perfect classification. When AUC persistently hovers around 0.46–0.54, it indicates that the model's confidence scores have virtually no monotonic correlation with the true labels — meaning the model has failed to extract any useful discriminative representation from the data.
What the Author Already Tried
To the author's credit, the experimental design was rigorous — well beyond the standard for most student projects. Approaches attempted included:
- EfficientNet-B0 for image-level binary classification
- BERT for analyzing OCR-extracted receipt text
- Numerical consistency checks (logical relationships between quantities, unit prices, subtotals, taxes, totals, and change)
- Multimodal fusion combining image, text, and numerical features
- YOLO for localizing tampered regions
- Five-signal ensemble: visual, textual, numerical anomaly, image artifact, and OCR rule signals
- Synthetic forgery data: generating fake samples by modifying amounts in real receipts
- DocTamper transfer learning and patch-level experiments using RGB/SRM features
The evaluation methodology was equally thorough: stratified five-fold cross-validation, PCA, feature auditing, and verification of annotation bounding-box round-trip consistency to confirm no data leakage between train/validation/test splits. When considering adding T-SROIE as supplementary data, the author specifically performed fold-aware deduplication since both datasets derive from SROIE and share source receipt overlap — yet this yielded no meaningful improvement.
This willingness to honestly report near-random results rather than chasing impressive-looking metrics with bigger models is precisely the hallmark of genuine research integrity.
SRM (Spatial Rich Model) is a classic feature extraction framework in digital image forensics, originally designed for steganalysis. It applies a set of high-pass filters to suppress the semantic content of an image and extract residual-level statistical features, making tampering artifacts such as compression traces, resampling signatures, and local noise inconsistencies visible. In patch-level experiments, combining SRM features with RGB appearance features in parallel is a common baseline in image forensics. DocTamper is a deep learning model specifically designed for document tampering detection, pre-trained on a large-scale document forgery dataset with some capacity for zero-shot cross-dataset transfer. However, a zero-shot AUC of approximately 0.50 indicates a significant domain shift between the tampering patterns DocTamper encountered during training and those in the Find It Again! dataset — the pre-trained knowledge failed to transfer effectively. This phenomenon is common in document forensics, as receipts from different sources vary greatly in scanning hardware, compression parameters, font rendering, and background texture.
Diagnosis: Data Problem or Methodology Problem?
The bootstrap confidence intervals for all current results include 0.50, meaning the models have not learned any effective discriminative signal. Adding 1,200 synthetic forgery samples produced no improvement; DocTamper zero-shot transfer sat around 0.50. These results point to a core conclusion: the problem likely exists at both the modeling and data levels simultaneously.
A common reason for synthetic data failure is that the model learns "synthetic-specific artifacts" rather than genuine forgery features. When you modify amounts using Photoshop-style operations, you typically introduce compression artifacts, font rendering differences, or edge discontinuities — artificial traces that don't exist in high-quality real-world forgeries. This causes the model to overfit on the synthetic set while achieving near-zero signal on the real test set.
Directions Worth Prioritizing for Validation
Addressing the six questions raised in the original post, here are recommendations oriented toward "verifiable hypotheses" rather than "swap in a bigger model."
Reframe the Problem: From Classification to Anomaly Detection
Receipt-level binary classification is likely the wrong primary objective. Since tampering affects only tiny regions, it's more principled to model this as a patch-level anomaly detection or segmentation task. Train a reconstruction or self-supervised model on authentic receipts to learn the distribution of "normal receipts," then flag regions with high reconstruction error as suspicious during inference. This sidesteps the severe class imbalance problem and better matches the localized nature of tampering signals.
Self-supervised anomaly detection works by training a generative or reconstruction model exclusively on normal samples to learn the distribution of normal data. During inference, regions that deviate from this distribution (i.e., with high reconstruction error) are flagged as anomalies. Common frameworks include autoencoder-based reconstruction error, normalizing flow-based log-likelihood estimation, and masked image modeling approaches (e.g., MAE). In document image scenarios, the model's learned "normal receipt" prior encompasses properties like consistent character spacing, font pixel density distribution, and ink noise patterns. Tampered digits often deviate statistically along these dimensions, even if visually the difference appears to be just a subtle digit change. Compared to fully supervised binary classification, this approach requires no large-scale labeled forgery samples and naturally handles severe class imbalance — making it a highly promising modeling paradigm for small-sample forensics tasks.
Representation Learning: Self-Supervised Pre-Training
With labeled data this limited, contrastive learning or masked reconstruction-style self-supervised pre-training is genuinely worth attempting. Pre-train on a large corpus of unlabeled receipt images so the model learns document-level priors around fonts, layouts, and textures, then fine-tune with the small labeled set. This is far more likely to converge on effective signals than training a classifier from scratch.
Numerical Consistency as a Standalone Strong Signal
The original post treats numerical validation as one component of an ensemble, but in practice, numerical logic contradictions are often the most reliable evidence of tampering — if subtotal plus tax doesn't equal total, that's a hard constraint, not a statistical feature. Evaluate the numerical consistency module's discriminative power in isolation; it may be the only component in the entire system that genuinely beats random.
Use Controlled Experiments to Distinguish Data Insufficiency from Method Failure
To answer "is this a data problem or a representation/generalization problem," design a diagnostic experiment: artificially inject highly visible, large-scale forgeries (far larger than real tampering) and check whether the model can learn to detect them. If it can't detect even obvious forgeries, the issue is with modeling or pipeline design. If it detects large forgeries but fails on subtle ones, this confirms that signal sparsity is the true bottleneck.
This type of diagnostic ablation is known in machine learning research as a "sanity check" or "feasibility probe" — its goal is not to directly improve performance, but to precisely locate the root cause of failure. Gradually scaling tampering from tiny (a single digit) to larger regions (entire rows or blocks) allows you to plot how detection ability varies with tampering salience. If the curve remains flat even for large-scale forgeries, the pipeline has a systemic flaw (e.g., incorrect feature extraction design or label alignment issues). If the curve rises sharply for large-scale forgeries, the method itself is sound and the bottleneck lies in the inherent weakness of real-world tampering signals. The latter conclusion also carries significant research value: it suggests the task itself may have a theoretical information-theoretic ceiling under current data conditions, providing a well-grounded anchor for expectations in future work.
Lessons for Small-Sample Forensics Researchers
The value of this case study isn't whether it ultimately produces a clean AUC score, but that it fully demonstrates an honest research process: exhausting mainstream approaches, rigorously auditing for data leakage, using confidence intervals to bound conclusions, and resisting the urge to throw compute at the problem.
For similar document image forensics tasks, a few lessons are worth keeping in mind: global classification is ill-suited for local tampering — anomaly detection and segmentation better fit the task's nature; synthetic data must be carefully guarded against shortcut artifacts; on small data, self-supervised pre-training and domain priors often outperform larger models; and most importantly, use falsifiable experiments to narrow down root causes rather than swapping models on intuition. Near-random results are themselves a valuable finding — they may faithfully reflect the objective limitations of currently available public datasets.
Related articles

Rysh Forge in Action: One OpenAPI Spec, Automatically Turned into Claude-Callable Agent Tools
Rysh Forge converts an OpenAPI spec into Claude-callable Agent tools, an MCP server, a Python SDK, and docs with one command — with built-in write confirmation and full audit logging.

OpenAI Agents SDK in Practice: Implementing Human-in-the-Loop Approval
A complete guide to implementing Human-in-the-Loop approval with OpenAI Agents SDK: pause tool calls with needs_approval, capture interruptions, handle approve/reject decisions, and serialize RunState for async workflows.

NeurIPS SAC Ticket Transfer Policy Sparks Debate: What's Behind the Change
NeurIPS SAC ticket transfer policy sparks debate in the Reddit ML community. We break down the controversy, likely reasons behind the change, and practical advice for researchers navigating top conference ticketing.