Unlocking Reliable GenAI: Strategies for Assessing LLMs in Real-World Applications
Evaluating LLMs is no longer an academic concern. As models get more intelligent and incorporate more modalities, gaining confidence in your application is only going to get harder. To get ahead early, the broader community needs to discuss practical solutions to calculate LLM performance & reliability across many metrics. In this talk, I will begin by providing a survey of the current approaches to evaluating LLMs and discuss their primary drawbacks - being too slow, expensive or biased. We will discuss practical solutions that will unlock faster iteration & more safety in GenAI, such as using tiny evaluators in an online setting & making efficient use of human feedback offline. At the end, you should be well equipped to understand what you can do today to get a clear signal on your GenAI application performance.