Project

LLM Evals

Medium
18 completions
~ 10 hours
4.0

By the end of this project, you'll build a complete evaluation pipeline for an LLM application. You'll gain hands-on experience with evaluation techniques such as analytics, human-as-a-judge, and LLM-as-a-judge. You’ll also learn how to use tools like Langfuse and Ragas to supercharge LLM evaluation. This project will help you ensure that your AI offers accurate recommendations and consistently meets high performance and reliability standards.

Provided by

JetBrains Academy JetBrains Academy

About

LLM evaluation is at the core of building trustworthy AI. In this project, you’ll work on a chatbot for a smartphone sales site, but the real focus is on assessing its performance. You'll use tools such as Langfuse and Ragas and various strategies to see how well the model delivers recommendations and comparisons.

Graduate project icon

Graduate project

This project covers the core topics of the Introduction to AI Engineering with Python course, making it sufficiently challenging to be a proud addition to your portfolio.

At least one graduate project is required to complete the course.

What you'll learn

Once you choose a project, we'll provide you with a study plan that includes all the necessary topics from your course to get it built. Here’s what awaits you:
Collect traces to analyze how the app is performing.
Tag, version control, and track metrics for prompts.
Use the Langfuse UI to annotate traces. Collect feedback and use it to score traces.
Set up evaluators for model-based evaluation.

Reviews

Dakouri Maurille-Constant Kobri
4 months ago
I've learnt how to perform an LLM evaluation implementing huma-as-a-judge, through feedback and manual annotation, and LLM-as-a-judge, using the Ragas framework. I enjoyed the experience of adding LLM evaluation tooling to my tech skills.
Shashank Gupta avatar
Shashank Gupta
9 months ago
It was a fantastic project to delve into the observability and evaluation of Agents. I learned how to integrate LangFuse's observability and monitoring framework into the agents, add user feedback per chat session, and manually annotate the agent's response after the chat. The best and most exciting ...
Brian Smith avatar
Brian Smith
1 year ago
This course prepares one to begin using LangFuse and understand why it is useful.

4.0

Learners who completed this project within the Introduction to AI Engineering with Python course rated it as follows:
Usefulness
4.6
Fun
3.8
Clarity
3.6