LLM evaluation defines test inputs, desired behavior, and scoring criteria to compare candidate outputs. Separating correctness, grounding, and format adherence helps reveal which changes actually help.
The demo uses illustrative scores to show candidates winning on different criteria. Check whether the dataset represents real tasks and whether grading is consistent.
When to use
Use it to compare model or prompt changes on repeatable examples.