28 min . Aug 3, 2026 . Education & Teaching
In this episode, we sit down with Dr. Jessica Cervi, a Learning Facilitator and Subject Matter Expert on AI for MIT xPRO. She'll walk us through harness engineering, an approach focused on building the environment an AI agent operates inside rather than just the instructions you hand it. Jessica breaks down how harness engineering differs from prompt and context engineering, why guardrails and evaluation are the two load-bearing pillars of the whole thing, and how using a more powerful model as a judge can catch the failures you would otherwise ship straight to your learners. We also dig into the tradeoff every builder runs into eventually: how much freedom to give an agent, and when the stakes of a task should pull you back toward tighter control.
Mentioned Links:
**Dr. Jessica Cervi (00:00)**
LLMs, and AI agents in particular, can be a little bit unpredictable, as we know. We ask them things, and even if our prompt is good and our context is good, sometimes they have a mind of their own. And so a term that has been trending in the world of AI is harness engineering, which is essentially the practice, or the set of practices, of creating a robust environment where your agent can operate. It lets you guide your agent and your workflows from point A to point B in a controlled manner.
**Luke (00:42)**
Hey folks, and welcome on in to Teach Us Something New. I'm your host. My name is Dr. Luke Hobson and I'm the Head of Instructional Design at MIT xPRO. This show is built around a simple idea: what happens when you sit down with some of the very best and brightest minds at MIT and ask them to teach you something new? That's exactly what we're going to be finding out together. In each episode, we'll sit down with a member of the MIT community and dive deep into their expertise, their research, and the ideas that drive them. Whether you are a lifelong learner, an educator, or you are just genuinely curious, you are in the right place.
Today we'll be sitting down with Dr. Jessica Cervi. She's a learning facilitator and subject matter expert for MIT xPRO and is an AI adoption partner for eDreams ODIGEO. In this episode, we are going to be learning about harness engineering.
What is harness engineering? That is literally the first question I asked Jessica, because I had never heard of the term before. She and I were talking the other day about her coming on the podcast, and I asked what we could talk about, and she said, "Have you covered harness engineering yet?" I said no, and that I had no idea what that even means. So we're going to be talking about this approach. We're also going to be diving into large language models and why they have a very difficult time evaluating their own output.
How do we actually get around that? And if we're going to do harness engineering, what types of guardrails do we need to put into place? We also get into AI agents and how to think about their controls, as far as how much or how little we should be doing with them. And we talk about what would happen if Jessica and I implemented harness engineering for our courses at MIT xPRO. What would we do differently? What is the first idea that comes to mind? So we get into how to actually use this information in the real world.
With all that being said, here is my conversation with Dr. Jessica Cervi.
---
**Luke (02:49)**
Jessica, welcome to the podcast.
**Dr. Jessica Cervi (02:51)**
Hey Luke, how are you?
**Luke (02:52)**
I'm doing well, I'm doing well. I'm excited to talk with you today because I don't know anything about today's topic.
**Dr. Jessica Cervi (02:57)**
Me too.
**Luke (02:59)**
So we are really going to be learning something new, which is the title of this show. But before I get ahead of myself, Jessica, would you mind introducing yourself to the audience and telling us more about who you are and what you do?
**Dr. Jessica Cervi (03:12)**
Of course. Hello everybody. My name is Jessica Cervi. A little bit about myself: I did my PhD in applied mathematics at the University of Saskatchewan. After that, I worked with various institutions as a subject matter expert and learning facilitator. One of those institutions was MIT xPRO, which is where I met you, Luke. And now my role is AI adoption partner at eDreams ODIGEO. I'm really, really happy to be here and connect again with you, Luke.
**Luke (03:45)**
Absolutely. So how did you actually get to MIT, Jessica? What is your origin story of how we ended up working together? It feels like forever ago that I've known you.
**Dr. Jessica Cervi (03:54)**
It feels like it, yes. So one of your colleagues, Indi Williams, was working at Emeritus at some point years ago, and Indi and I worked on a project together. It was an AI course as well. We were designing it together. And then when she moved to MIT, she contacted me again, we reconnected, and I started working with MIT xPRO. So, full circle moment. It was really nice.
**Luke (04:25)**
Well, we've been very lucky to have you as a subject matter expert and learning facilitator for a number of different AI courses and programs. You obviously have a background in AI and you've helped design many of the courses we've been talking about. And I know that you are currently specializing in harness engineering, because you told me before coming on the show. For someone who has never heard of this term before, like myself, what is harness engineering?
**Dr. Jessica Cervi (04:54)**
Before we introduce harness engineering, let's start with two terms the audience has perhaps heard of. The first one is prompt engineering. Prompt engineering is the art of designing a prompt for your LLM or your AI agent so that the responses are more tailored to you, more specific.
As AI capabilities evolved, we realized that prompt engineering was not enough. And so we moved into something called context engineering. This is the practice of essentially feeding your LLM or your AI agent with a memory, with documents. Think of it that way.
Now, LLMs and AI agents in particular can be a little bit unpredictable, as we know. We ask them things, and even if our prompt is good and our context is good, sometimes they have a mind of their own. And so a term that has been trending in the world of AI is harness engineering, which is essentially the practice, or the set of practices, of creating a robust environment where your agent can operate. It lets you guide your agent and your workflows from point A to point B in a controlled manner. If that makes sense.
**Luke (06:29)**
All right, so I'm following along with this. You just mentioned something I wanted to ask about. When I searched for the term harness engineering, I saw one of the most popular posts on Reddit saying that prompt engineering was the big thing in 2024, context engineering was the big thing in 2025, and asking whether harness engineering is the big thing in 2026 and 2027.
**Dr. Jessica Cervi (06:53)**
Potentially, potentially. As AI agents are becoming, well, they're probably the biggest thing in 2026 so far, right? I think that the practice of controlling them, which is harness engineering, is just going to become more and more popular. So becoming familiar with this term right now can put you in a comfortable position for the future.
**Luke (07:18)**
How did you end up specializing in this?
**Dr. Jessica Cervi (07:23)**
To be honest, just trying to stay at pace with AI. We started with traditional AI and deep learning and all of that, and then when ChatGPT launched a few years ago, I kind of got addicted to it in a good way. From there, I try to stay at pace as best as possible. Obviously it's a rapidly evolving environment, so it's difficult to know everything all the time, but I love it, and that's where I'm focusing my career.
**Luke (07:57)**
How are you currently staying on top of trends within the AI space, by the way? I think that's the number one thing most people are trying to figure out. As this keeps shifting and changing, it's really tough to keep up. I know for myself, I have essentially trained my LinkedIn algorithm to show me what all my smart friends and colleagues are saying about what we should be looking out for in the future, so then I start to experiment and dabble with everything. But do you have any techniques for staying up to date?
**Dr. Jessica Cervi (08:32)**
Yeah. My obvious cheat code is that AI is my job, so a hundred percent of my time is spent using AI. For me, perhaps it is a little bit easier to stay on top of it. But one trick that I use, and I think a lot of my colleagues do as well, is to have an AI agent that scans the articles or podcasts or YouTube videos that they like the most and that are most relevant. Once or twice a week, they just get a report with everything that's becoming relevant and important. You sort of get your own personalized newsletter. You still have to read and to practice, but it can help you navigate the insane amount of information that we're getting these days.
**Luke (09:27)**
I love that. That's a fantastic tip. That was actually one of the first things I did when custom GPTs became very popular. It was literally that: how do I stay on top of the latest and greatest with learning science? There are a thousand different articles and journals and everything else. And as you were saying, that became my little personalized newsletter that said, okay, here's the latest for today. So that makes a lot of sense.
Let's go back to harness engineering before I get us too far off track. First, is this the correct terminology? If I say I opened up a harness inside of an LLM, is that a correct way of saying it, or no?
**Dr. Jessica Cervi (10:08)**
Not really. It's not a harness per se. There can be files that make up the harness, but the harness itself is not a file. It's not as simple as configuring some variables for your agent. There are some practices, some pillars, that define harness engineering, and essentially implementing these five pillars can help you build a robust harness.
**Luke (10:37)**
Okay, so knowing the correct terminology, if I'm looking under the hood and trying to see what is inside, what am I looking at? Is it sophisticated code? Is it a series of prompts, instructions? How is this being built?
**Dr. Jessica Cervi (10:55)**
In a few different ways. Prompts or markdown files are one component of a harness. Then there are things like guardrails. Guardrails are essentially pieces of text that you embed in your prompt or in your agent skills that tell the agent what not to do. That is as simple as plain text, but it is part of the harness.
It can also be code. If you want to hard code something in your agent, something deterministic that you always want the agent to do, that can be code.
And perhaps the biggest part of harness engineering is the evaluation. As we work with an agent and as we modify the agent, how do we make sure that the version we are using today is better or worse than the one we were using last week? Building evaluations for agents is extremely important. It's very overlooked, but it is extremely important. And that can be as simple as some code, or a table, or a rubric. That is all part of the harness as well. So it's a mix, in short.
**Luke (12:13)**
It's a mix. You mentioned briefly that there are five different pillars. Is that part of the pillars?
**Dr. Jessica Cervi (12:16)**
Yes. The evaluation suite is one of the pillars. Being able to implement observability in your agent and keep an eye on the performance, and whether the performance is good or degrading, is extremely important, especially if you have agents that interact with real users. That's part of the harness as well.
**Luke (12:41)**
That makes sense. Sticking on the topic of evaluation, when I ask people to dive in further on taking an output and verifying whether the information is correct and reliable, one trick I've heard a lot of people use is to have one LLM check the work of another LLM. Of course there is human intervention in the loop somewhere as well. But is this even a viable method, or is it just not as great as you think?
**Dr. Jessica Cervi (13:09)**
No, no, it's actually really good. This method is called LLM as a judge. Essentially you're using another LLM, typically more powerful than the one that powers your agent, to check the agent's work. The way that LLM as a judge operates is that it's given a rubric. And as educators, you and I know very well what a rubric consists of. You create this rubric with a grading system for your LLM as a judge, and you define inside the rubric what a good response looks like and what a medium response looks like. Then the judge is able to use that rubric to grade and to judge the response. It's a widely used method.
If I can add one more thing on LLM as a judge: it's better to have specialized judges, so as not to overwhelm them. Keep your rubrics small and use multiple judges for different tasks.
**Luke (14:17)**
So why can't we give that type of rubric to our original LLM and say, I want you to evaluate your own work? Why is that not the best idea?
**Dr. Jessica Cervi (14:26)**
Because that would be like cheating. It's like giving a student a rubric and asking them to grade their own work. Of course it's going to score perfectly, because it thinks its own response is the best. But once you ask a different LLM to judge that with a rubric, chances are you're going to get a more fair response.
**Luke (14:49)**
That's interesting. Is there a way to design around that? The reason I'm asking is that I know some institutions only have a license to work with one LLM. Dartmouth, for example, has a partnership with Anthropic. So if we're saying it's great to use Claude for this, but what about using another LLM as the judge, is it possible to still use a single provider, or is it always wise to find multiple?
**Dr. Jessica Cervi (15:26)**
The good news is that Anthropic, and other providers as well, come with multiple models. Anthropic has three tiers of models: Haiku, Sonnet, and Opus. Typically within OpenAI or Google or Claude, there is always a tier of model you can choose from.
It is better to have the LLM as a judge be powered by a more powerful model, but if there's no way around it, using an equivalent model is also acceptable. The LLM as a judge is always going to be only as good as the rubric that you define. If the rubric is poor and imprecise, then the LLM as a judge has nothing to refer to.
**Luke (16:17)**
That makes sense. I didn't even think a model would be that effective when you use another model made by the same organization.
**Dr. Jessica Cervi (16:26)**
It actually works quite well. It's a very smart way of using AI to check AI's work, because humans would not be able to check all of it. With the amount of data and code that AI produces, we just don't have the capacity to review all of that.
**Luke (16:44)**
No. And I know so many people who wonder about using a free trial of one tool versus a paid version of another, and what that looks like. Just from using many different free ones over the years, it's just not optimal. Hallucination risk goes up, and everything else of the sort.
So that's fascinating. How do you decide how much control to give to that model versus how much structure to impose? I'm assuming that if you give too much control, it's going to go one way, and on the other hand, if you tighten it too much, then it can't get the job done. So how do you know what to do for that balance?
**Dr. Jessica Cervi (17:24)**
A couple of things. For low stakes decisions, you can probably give it a little bit more freedom. For example, if it's an internal agent, not customer facing or anything like that, it's fine to be a little bit more loose and play with it.
But if your agent has to process things like payments or sensitive information, in that case your harness has to be very robust. I recently read a paper by OpenAI, a white paper or an article, and essentially the key takeaway was that the best agents are the ones that have a robust harness, so they know what to do in every situation. Otherwise you just risk running into infinite loops of nonsense, and it becomes a bit difficult to get the response that you want.
**Luke (18:28)**
That makes sense. I'll try to find the article and link it in the show notes.
**Dr. Jessica Cervi (18:31)**
I'll send it to you.
**Luke (18:33)**
Thank you. As you were talking, I realized there's one thing I haven't asked you about yet. Can you give us an example of this in action? What does it actually look like, and how could it be useful for someone listening to the show right now who wants to go and dabble in harness engineering?
**Dr. Jessica Cervi (18:50)**
Yes. The easiest way to think about it is with the tools. An agent connects to the outside world through connectors, through MCPs, and those are essentially the gateway for the agent to connect with different apps or software. What happens is that if you allow the agent to use one MCP to its full ability, you may risk the agent performing actions that are dangerous or harmful. For example, the BigQuery MCP can potentially have the ability to delete a database or delete some rows in the database.
If you control what the agent can do with that MCP, that is harness engineering. You're telling the agent which functions in that MCP it is allowed to perform, so that you don't need to worry. You know that your database and your rows are safe. It can be as simple as deciding how much liberty an agent has with a particular MCP. That's a very simple practice of harness engineering, but it falls under that as well.
**Luke (20:22)**
So hypothetically speaking, if we could use this technology for our online courses, what would you try to do? In a perfect world.
**Dr. Jessica Cervi (20:31)**
Ooh. Well, between you and me, although people will listen to this, I would probably try to implement an LLM as a judge to grade all the assignments.
**Luke (20:46)**
I mean, that is fair. There are a number of institutions right now that truthfully are trying to figure out how to give feedback in an appropriate way for different types of assignments, especially if you have courses with hundreds upon hundreds of students and only one faculty member. What do you do?
**Dr. Jessica Cervi (21:05)**
Exactly.
**Luke (21:05)**
I actually saw a demo of an example of this many years ago. It was before the pandemic, because we were still in the office as normal. I remember at the time I thought there was no way this was good enough. And then they explained more, and said that if you train it with enough data and tell it what to say and what to do, it works. Now here we are several years later, and I'm seeing companies marketing this and saying they can help with personalized learning, with different sources of feedback. And I'm thinking, this is now a thing. This is the future of that.
**Dr. Jessica Cervi (21:45)**
It became a thing very quickly.
**Luke (21:45)**
It's fascinating. Of course, when you go down that line of thinking, you have to consider it not just from the technology perspective but from the human perspective. Someone might say, no, I want a person's feedback rather than the automated version. Or maybe they can figure that out for themselves because they've already done that before, submitting work for that feedback, which some do.
How do you decide about the outputs you're receiving? You mentioned having a rubric to say that something is good, that it's effective. But do you go through a set of trials, a pilot test? Do you do anything along those lines to test whether this works as intended?
**Dr. Jessica Cervi (22:37)**
Yes. Before even releasing any agent into production, running multiple evaluation suites is extremely important, because that can give you a benchmark not just on accuracy but also on the quality of the responses.
What's even more important is that this monitoring is kept alive even when the agent is in production, because we don't know how a user will interact with an agent. We can try to predict that when we build the agent, and we can build our evaluation suite around that, but we really don't know what kinds of questions people are going to ask. So continuously looking at that data and testing the agent even while it's in production is important for that reason.
**Luke (23:31)**
That makes sense. Speaking of people using agents for the wildest random reasons, did you see the trend going around where people were using the agents on the Chipotle and McDonald's websites to help them out with their code?
**Dr. Jessica Cervi (23:45)**
Exactly. I was actually thinking of that example. It was something. It was like, I need help with my Python homework.
**Luke (23:52)**
Right. Might as well just go to McDonalds.com and figure out if it's going to help me. And it's funny, because it was actually giving answers.
**Dr. Jessica Cervi (24:00)**
And that's exactly it. Nobody put the guardrail inside that chatbot saying, do not give any information that's not in our company policy. It could have been a little bit more complicated than that, but that was the main idea.
**Luke (24:15)**
Yes, that makes sense. With evaluation methods, I wanted to ask about the tradeoff between how elaborate a harness is and what it costs to run. I'm thinking about time, money, complexity. Is there a tradeoff there?
**Dr. Jessica Cervi (24:33)**
There is a tradeoff. Obviously, investing in a harness is time consuming. Agents are difficult to build, and robust agents are even harder to build than a lot of people think. So obviously the time investment is there.
Cost-wise, MCPs can be expensive if not toggled correctly, and so the harness can actually help you save a little bit of money. Evals can be a little bit costly, but if they're well designed, we're talking dozens of dollars, not hundreds or thousands. Nothing of that sort.
Ultimately, the time investment at the beginning, when designing an agent, is paid off. If you have a robust agent, the time investment you spend at the beginning, you're going to get back. Because you want a robust agent. You don't want one that you design quickly and that fails even quicker. You want one where you spend a little bit of time designing it and it works.
**Luke (25:51)**
That makes sense. Is this here to stay? Do you think this is what we're going to be talking about for the long haul, or is a future model going to come out that changes everything?
**Dr. Jessica Cervi (26:04)**
Hard to say. I think that for maybe the next six months we are safe. But if any model comes out that has some auto-evaluation ability, that would be almost groundbreaking, because evaluation is the hardest part. The LLM output is probabilistic, and so we can almost never predict what the LLM is going to tell us. Unless something changes in that sense, I think harness engineering, and evals in particular, are going to be relevant for a while.
**Luke (26:45)**
That's good to know. Until something new comes out and we all have to learn a brand new thing.
**Dr. Jessica Cervi (26:48)**
Exactly.
**Luke (26:50)**
Nonstop learning in the world of AI is what I've learned over the last couple of years. Well, Jessica, thank you so much for coming on the show and talking more about harness engineering and your background. Where can people go to learn more about you and your work? Where can they find you?
**Dr. Jessica Cervi (27:05)**
You can find me on LinkedIn. I'm not super active, but we can connect there if they're interested, for sure.
**Luke (27:11)**
Awesome. I will be sure to put those links in the show notes as well. Once again, Jessica, thank you so much for coming on the show. I appreciate it.
**Dr. Jessica Cervi (27:17)**
Thank you so much, Luke. Have a good day.
**Luke (27:18)**
Folks, I hope you enjoyed my conversation with Dr. Jessica Cervi. If you did, be sure to subscribe to the podcast and give it a five-star rating wherever you are listening. More episodes are coming out each month, where we talk with different members of the MIT community. So if you enjoyed this one, certainly stay tuned for more.
And if you are looking to learn more about AI and how to use it in your workplace, you can check out some of the courses that Jessica has helped us design by going to the show notes or to xpro.mit.edu.
Other than that, folks, that is all I have for you. Stay curious, keep on learning, and I'll see you next time.