LLM evaluation

LLM 평가

Compare model or prompt quality on representative tasks using explicit criteria.

···
html
<div class="scene"><div class="heading">EVAL MATRIX · EXAMPLE</div><div class="matrix"><div>후보</div><div>정확성</div><div>형식</div><div>A</div><div id="a-accuracy" class="score">높음</div><div id="a-format" class="score">낮음</div><div>B</div><div id="b-accuracy" class="score">보통</div><div id="b-format" class="score">높음</div></div><div class="muted" id="eval-note">기준을 선택해 차이를 확인</div></div>
css
.scene{width:min(94vw,800px);height:min(88vh,326px);padding:clamp(10px,2.7vmin,19px);border:1px solid var(--line);border-radius:13px;background:var(--surface);font:500 clamp(13px,2.4vmin,17px)/1.3 var(--font-sans,sans-serif);position:relative;overflow:hidden}.scene .mono{font-family:ui-monospace,SFMono-Regular,monospace}.scene .muted{color:var(--muted)}.scene .accent{color:var(--accent)}.scene .heading{font-weight:750;color:var(--accent);margin-bottom:clamp(5px,1.6vmin,12px)}.matrix{display:grid;grid-template-columns:repeat(3,1fr);gap:3px;height:65%}.matrix div{display:grid;place-items:center;text-align:center;border:1px solid var(--line);border-radius:4px;background:var(--bg)}.matrix .score.active{background:color-mix(in srgb,var(--accent) 20%,var(--surface));border-color:var(--accent);color:var(--accent)}.matrix div:nth-child(-n+3){font-weight:700}.muted{margin-top:7px}
js
let format=false;function draw(){format=!format;document.querySelectorAll('.score').forEach(x=>x.classList.remove('active'));document.getElementById(format?'b-format':'a-accuracy').classList.add('active');document.getElementById('eval-note').textContent=format?'형식 기준: B 우세':'정확성 기준: A 우세'}draw();setInterval(draw,1450)

LLM evaluation defines test inputs, desired behavior, and scoring criteria to compare candidate outputs. Separating correctness, grounding, and format adherence helps reveal which changes actually help.

The demo uses illustrative scores to show candidates winning on different criteria. Check whether the dataset represents real tasks and whether grading is consistent.

When to use

Use it to compare model or prompt changes on repeatable examples.

Open as page ↗