A generative model composes likely text, so fluency is not proof of truth. It may misread supplied evidence or fill gaps with unsupported claims.
The demo compares a date in evidence with a date in an answer and marks the mismatch. String matching cannot detect every factual error; important claims need independent checks.
When to use
Account for it when reviewing model answers or defining grounded-answer quality.