LLM 평가

LLM evaluation

대표 과제와 채점 기준으로 모델·프롬프트의 답변 품질을 비교합니다.

···
html
<div class="scene"><div class="heading">EVAL MATRIX · EXAMPLE</div><div class="matrix"><div>후보</div><div>정확성</div><div>형식</div><div>A</div><div id="a-accuracy" class="score">높음</div><div id="a-format" class="score">낮음</div><div>B</div><div id="b-accuracy" class="score">보통</div><div id="b-format" class="score">높음</div></div><div class="muted" id="eval-note">기준을 선택해 차이를 확인</div></div>
css
.scene{width:min(94vw,800px);height:min(88vh,326px);padding:clamp(10px,2.7vmin,19px);border:1px solid var(--line);border-radius:13px;background:var(--surface);font:500 clamp(13px,2.4vmin,17px)/1.3 var(--font-sans,sans-serif);position:relative;overflow:hidden}.scene .mono{font-family:ui-monospace,SFMono-Regular,monospace}.scene .muted{color:var(--muted)}.scene .accent{color:var(--accent)}.scene .heading{font-weight:750;color:var(--accent);margin-bottom:clamp(5px,1.6vmin,12px)}.matrix{display:grid;grid-template-columns:repeat(3,1fr);gap:3px;height:65%}.matrix div{display:grid;place-items:center;text-align:center;border:1px solid var(--line);border-radius:4px;background:var(--bg)}.matrix .score.active{background:color-mix(in srgb,var(--accent) 20%,var(--surface));border-color:var(--accent);color:var(--accent)}.matrix div:nth-child(-n+3){font-weight:700}.muted{margin-top:7px}
js
let format=false;function draw(){format=!format;document.querySelectorAll('.score').forEach(x=>x.classList.remove('active'));document.getElementById(format?'b-format':'a-accuracy').classList.add('active');document.getElementById('eval-note').textContent=format?'형식 기준: B 우세':'정확성 기준: A 우세'}draw();setInterval(draw,1450)

LLM 평가는 입력 사례, 기대 동작, 채점 기준을 정하고 후보 답변을 비교하는 과정입니다. 정확성, 근거성, 형식 준수처럼 서로 다른 기준을 분리하면 어떤 개선이 실제로 도움이 되는지 볼 수 있습니다.

데모의 점수는 가상의 예시이며 두 답변이 기준별로 다른 결과를 받는 장면을 보여줍니다. 평가셋이 실제 사용자 과제를 대표하는지와 채점의 일관성도 함께 확인해야 합니다.

언제 쓰나

모델 교체나 프롬프트 수정 전후를 감각이 아닌 재현 가능한 사례로 비교할 때 사용합니다.

페이지로 열기 ↗