Artificial Intelligence · LLMs
What is a common mistake when working with LLM evaluation?
Using only subjective manual testing or a single aggregate score can hide important failure modes. In practice, the value of this concept comes from understanding where it belongs in the architecture, what problem it solves, and what constraints or tradeoffs it introduces.