Portrait of Zahra Anvari
Machine Learning Researcher & Engineer

Zahra Anvari

Document AI · Large Language Models · Multimodal AI · Robust Machine Learning

I am a machine learning researcher and engineer with a Ph.D. in Computer Science from the University of Texas at Arlington. My work focuses on large language models, Document AI, structured information extraction, multimodal learning, and evaluation of robust AI systems under realistic data-quality constraints.

About

My research combines academic investigation with industry experience building production-scale machine learning systems. I previously worked as a Senior Machine Learning Engineer at Iron Mountain, where I developed Document AI pipelines involving document enhancement, OCR, transformer-based information extraction, table detection, and LLM-powered question answering.

I am currently conducting independent research on LLM-based key–value extraction from forms and receipts, with an emphasis on reproducible evaluation, OCR robustness, prompt sensitivity, and failure analysis. My broader goal is to build document understanding systems that remain reliable when deployed on noisy, visually complex, and imperfect real-world data.

Large Language Models Document AI Key–Value Extraction OCR Robustness Multimodal AI Model Evaluation Synthetic Data Information Extraction Retrieval-Augmented Generation

Recent News

2026
Completed From Pixels to Pairs, a controlled benchmark of LLM-driven key–value extraction under realistic OCR noise.
November 2025
Published a comprehensive survey of large language models for structured document understanding on TechRxiv.
2026
Contributed bug fixes and tests to open-source machine learning libraries, including Hugging Face Evaluate and Unstructured.

Featured Research

Current work in robust document understanding, structured extraction, and evaluation of large language models.

LLM Evaluation · Document AI · OCR Robustness

From Pixels to Pairs

A reproducible benchmark of six open-source instruction-tuned LLMs for key–value extraction across FUNSD, CORD, and SROIE under clean-text and OCR-degraded conditions. The study compares PaddleOCR, EasyOCR, and Tesseract and analyzes hallucination, numeric corruption, prompt sensitivity, and key–value misalignment.

Survey · LLMs · Multimodal Document Understanding

Large Language Models for Structured Document Understanding

A comprehensive survey and interpretive review of LLM-based and traditional approaches for key information extraction, named entity recognition, document classification, and document question answering, with emphasis on architecture, modality fusion, prompt sensitivity, and generalization.

Computer Vision · Image Restoration · GANs

Enhanced CycleGAN Dehazing Network

An unpaired image-to-image translation framework for single-image dehazing using global-local discrimination, perceptual supervision, color consistency, residual learning, and skip connections.

Benchmarking · Image Dehazing · Dataset Creation

Sun-Haze Benchmark

A benchmark and dataset for evaluating single-image dehazing methods under non-uniform, realistically colored sunlight haze using full-reference and no-reference image-quality metrics.

Publications

Selected publications. See Google Scholar for the complete and most current list.

2026

From Pixels to Pairs: A Comprehensive Benchmark of LLM-Driven Key–Value Extraction in Noisy Document Settings
Zahra Anvari and Vassilis Athitsos
Manuscript, 2026

2025

Large Language Models for Structured Document Understanding: A Comprehensive Survey of Tasks, Models, and Challenges
Zahra Anvari and Vassilis Athitsos
TechRxiv, 2025

2021

A Survey on Deep Learning Based Document Image Enhancement
Zahra Anvari and Vassilis Athitsos
arXiv preprint, 2021
Enhanced CycleGAN Dehazing Network
Zahra Anvari and Vassilis Athitsos
VISAPP, 2021

2020

Evaluating Single Image Dehazing Methods Under Realistic Sunlight Haze
Zahra Anvari and Vassilis Athitsos
ISVC, 2020

2019

A Pipeline for Automated Face Dataset Creation from Unlabeled Images
Zahra Anvari and Vassilis Athitsos
ACM PETRA, 2019

Experience

2024 – Present

Independent Researcher

Research on LLM-based Document AI, robust key–value extraction, OCR sensitivity, reproducible evaluation, and hybrid document understanding systems.

2021 – 2024

Senior Machine Learning Engineer · Iron Mountain

Developed production Document AI systems involving document restoration, OCR, transformer-based information extraction, table detection, RAG, and model evaluation.

2019

Deep Learning Intern · Wave Computing

Worked on face and emotion recognition, super-resolution, object tracking, and model benchmarking.

2016 – 2021

Graduate Research Assistant · University of Texas at Arlington

Conducted research in deep learning, computer vision, image dehazing, and automated dataset generation.

Teaching

I have several years of university teaching assistant (TA/GTA) experience across computer science courses, including machine learning, algorithms and data structures, parallel processing, computer architecture, software engineering, and computer systems. My teaching interests include machine learning, artificial intelligence, computer vision, data structures, and applied programming.