Accepted · ASONAM 2026 Industry Track

SemCorBench: Measuring Semantic Corruption in LLM-Derived Data Artifacts

A reproducible benchmark framework for measuring meaning-level failures that can pass conventional structural validation.

Research question

How can teams measure semantic corruption in LLM-generated and LLM-derived data artifacts when records remain structurally valid?

Contribution

The work frames semantic corruption as a measurable data-quality problem across generation outputs, retrieval-grounded responses, factual-verification labels, and ETL-style numeric records. It is designed to compare task-specific validators without treating schema validity as evidence of factual correctness.

Why it matters

LLM-derived artifacts can look operationally healthy while carrying unsupported labels, distorted values, or meaning-level errors into downstream analytics. A benchmark makes those failure modes easier to test, compare, and discuss.

Publication status

Accepted for the ASONAM 2026 Industry Track. The official manuscript link will be added after publication. This page intentionally provides only a high-level summary before proceedings publication.