Video: The Ultimate Guide to Testing AI and Agentforce in Salesforce | Duration: 3552s | Summary: The Ultimate Guide to Testing AI and Agentforce in Salesforce | Chapters: Introduction to Webinar (4.64s), Salesforce Ben Introduction (131.915s), AI Testing Fundamentals (215.895s), AI Cost Challenges (838.11005s), Testing AI Systems (965.66003s), Testing AI Agents (1102.265s), Benchmarks and Productivity (1211.6449s), Testing AI Tools (1535.0701s), Testing Agent Force (1894.06s), Wrapping Up AI (2418.265s), Monitoring AI Quality (3108.725s), Genetec AI Applications (3299.77s), Prioritizing AI Use Cases (3320.33s), Conclusion and Thanks (3488.125s)
Transcript for "The Ultimate Guide to Testing AI and Agentforce in Salesforce": Hello, everyone, and welcome to another Salesforce Ben webinar, this time in partnership with Tricentis. And we're gonna be talking about a super popular topic that I don't think has really been discussed that much in the ecosystem at the moment, but it's very important, in the in the age of AI and agent force, which is testing the the agents themselves. So, let us know where you're dialing in from in the chat, and we will get started in just a second. Just wait for a few more people to dial in. We're gonna be having a q and a at the end as well. So if you've got any questions throughout the session, then feel free to put them in the q and a box, and we'll try and get around to as many as possible. Hey, Jacqueline. Hey, Monica. Thanks for being with us today. Alright. Let's get started then. So welcome, everybody, again. We're going to be talking today about the ultimate guide to testing AI and Agent Force in Salesforce. So although Agent Force is still a pretty new product, it's becoming adopted very quickly. And sales also talks about testing a bit, but, you know, not in kind of a a holistic way. But when I was speaking to someone the other day who just completed an agent force implementation for a 60,000 person organization, He told me that testing requires about three times the amount of time than implementation does, which is probably quite unique compared to, normal Salesforce implementations. But it's the fact that agents could do anything. Right? It's not just populating a field or doing some basic automation. The the responses could be numerous, so you've got to worry about toxicity detection, hallucinations, and, generally, the agents just going a bit off rails. So that's what we're gonna be talking about today with Tricentis. Just to introduce Salesforce then very quickly, so we are the largest Salesforce media company in the ecosystem. You'll find a few of the things that we do on this slide at the moment. So, we publish about 20 new articles a week for a broad range of, disciplines such as business analysis, admins, developers, consultants. So no matter what you are in the ecosystem, we've got you covered. We've got our weekly newsletter that goes out. It's a fantastic way to, keep yourself informed about everything that's going on in the ecosystem, new releases, new features, news that's happening. We've got events such as this one, but also in person events. So if you're New York World Tour, or London World Tour, or Dreamforce, we're gonna be there with physical happy hours and, breakfast events. We've got our YouTube channel and also our podcast as well that you can subscribe to on your favorite channel. So let's introduce the speakers today. So just to introduce myself very quickly, my name is Ben McCarthy. I'm the founder of Sales with Ben, and I'm gonna be your host today. And I am joined by Julian and David, who I'm gonna hand over to you in a second. I will meet in a second, but I will be back at the end to facilitate the q and a. So as I said, if you've got any questions throughout the session, feel free to put them in the q and a chat box, and we'll get around to as many as possible. Julian and David, over to you guys. Yeah. Thanks, Ben. I'm Julian Ho, lead product manager here at Tricentis. Hi. I'm David Goel, the VP of AI and machine learning here at Tricentis. Alright. So I've been given the, unique honor of kicking this off. It's a bit of a difficult topic, how to test artificial intelligence in general. Like Ben said in the introduction, it's a bit of a weird animal. There's, times when it goes off the rails. It can be toxic. It sometimes feels a little bit like we've added a person to our environment, and we're thinking of it as a machine. And so we're gonna go through a little bit of some foundational stuff, like what is it that makes testing AI, distinct and unique? What are agents in this new language? What does it all mean? What are some of the unique benefits of testing, AI based applications? And then diving a little bit more into the sort of Salesforce and, AgenTek work within Salesforce and Agentforce. So, let's kick it off. Over here on the next slide, we're gonna do a brief introduction to what it is with, artificial intelligence and testing. So the first thing that we need to dive into is we've used a lot of terms. We've used terms like agent. But what does it actually mean? Well, let's start with a term that we're all actually pretty familiar with, and that's the term of automation. Now often we think, hey. An agent's a form of automation, but, not strictly true. An automation is something that is programmatically defined. It's deterministic. So think like you're creating an account. You expect the account to be created. You don't expect the machine to come back to you and say, I don't really wanna create an account. It's a it's a specific thing that is going to happen. Similarly, when you're using a calculator or using Excel and you say, calculate the standard deviation or sum this up, there's no question that you know exactly what the output is going to be. You don't necessarily know exactly how to get to that output, but you know exactly what that output is going to be. Or at least, maybe I'm only speaking for myself. It's been a while since I did the math behind calculating a standard deviation. But even though I don't know exactly what that formula is, I have great confidence in the underlying mechanism that it will do exactly what it says it's going to do. AI workflows are a little bit different, and we actually have a little bit of a a slide graphic underneath this to to demonstrate it. An AI workflow is where you say, well, I'm not any longer just going to consider straight up logic, like it's going to be a plus b equals c. There's a little bit of fuzzy in here. And we when we say a little bit of fuzzy, we're often talking about a little bit of AI. And this might be the case that, someone is coming in, they're gonna give me a request like, hey. I'd like to buy one of your products. And an AI workflow will be what takes them through that process. Now an AI workflow still has a lot of programmatic underpinnings. It's very linear. It kind of has one input, and then it expects one output at the end of it. The last thing that we're gonna talk about here is an AI agent, and this is kind of where we, surrender to the madness, if you will, and say, the whole thing is going to now be managed by an intelligent machine. We're going to try to put some guardrails around it, but we're really looking to say, we don't know the process that needs to be taken here. We just have a good idea about the type of thing that our customer needs to do, and we want an intelligent machine to help guide our customer, guide our staff, or complete some task on our behalf. So we've talked a bit about the core foundations, and it really is that gradation from pure logic to pure fuzziness. But there are some other things that we need to consider. Another thing that we need to consider is what type of task you would allocate to each one of these categories. We know what type of task we'd do with automation. We would give it predefined tasks, things that we know all of the steps involved. AI workflows, it's a little bit fuzzier. We might say, well, we expect some degree of flexibility from the machine, but we still it fits within a a defined lane. But for an AI agent, it starts to lean all the way into that fuzzy type of task. We're saying, we know generally what a good outcome looks like, but everything in between, like, even the types of requests we're gonna get, everything in between that and the outcome, we're going to allow to be adaptive. Another thing that we need to consider in this world is, like, what are the different various strengths and weaknesses of this? So I've actually got the strengths here. We can bring up the weaknesses as well. We can talk about them at the same time. So the strengths of automation are fairly obvious. You've got speed, reliability. It's just super fast to execute. You know exactly what you're getting. The limitations are that there's a huge amount of effort put into, building automation. So it's sort of like you pay for that effort upfront, and it's limited. It can only do the specific thing that you told it to do. AI workflows are much better at handling, the complex rules where you're kind of delegating a bit of that decision making to the algorithm, but they require data to train. They require models. They require benchmarking, and they can they're a little bit harder to debug and interpret. For example, if you had an AI workflow that was, gauging sentiment analysis and then sending an email to the people who had, a poor experience just as a follow-up, you wouldn't be able to tell necessarily why the fuzzy part, which is sort of if you draw this as three boxes, you have incoming, feedback, then you've got sentiment analysis box, and then you've got send email, which is a decision like positive sentiment. Send them an email saying, hey. We're glad you enjoyed our product. You know? Would you like to rate our product? And then negative sentiment would be like, hey. Would you mind if someone reached out to you understand what your negative sentiment is? So this part, you're confident in. Like, you know what the routing is. That's the Boolean logic part. But the part in the middle where it says what's the actual sentiment, you really have no idea why the AI made that decision. So you're starting to see that gradation of power and interrogability or, transparency trade off, And you bet heavily on that trade off when you go into AI agents because you're saying, right, I can give these new tasks. It's super rapid, super easy to get started with, but you're trading transparency. You don't really know what's going to happen when you release this. So you test and test and test in order to get varying degrees of confidence that this, fuzzy machine that you've put into your environment is actually going to produce the results that you want. The last thing that we want to talk about on this slide is a few examples. Now you've got, some simple ones like an algorithm to send a Slack notification every time a news, lead signs up. Great. Exactly known. Analyze and score a website inbound lead using ChatJPT. This is kind of that process where you go a, b, c, and the b part, which is the AI, is sort of the analysis and scoring part. An AI, agent is sort of, okay, some inbound lead comes in, and I want you to just go find out stuff about them. Tell me what they have, what they like, what they're doing. And that's where you're starting to see a huge amount of power. It would be incredibly hard to create an automation that does that, but, you don't really know what websites it's gonna visit. How is it gonna categorize that information? So that's a bit of, creating the foundation, if you will, about automations, workflows, and agents. So let's go over to the next slide where we're gonna talk talk a little bit more about some of the new problems that we've introduced with, the advent of generative AI and bringing aboard these AI agents. The first sort of new problem that we've introduced is the concept of self driving AI. This is where you're giving control over to the AI and you're allowing it to take actions with your hands off the wheel, if you will. There's a good analogy here in self driving cars where all of the AI agent vendors, and I think Salesforce does this as well, but groups like True dot ai especially, will say you're responsible for what the agent does. You're supposed to supervise it. In the same way that if you, drive a vehicle that's got a self driving mode, they'll often say, you're gonna drive yourself, but you have to stay behind the wheel. Sort of a liability protection. If the agent takes some kind of, erroneous or malicious action that your job is to correct it and prevent it. And that seems counterintuitive to us because we're like, well, if I have to supervise it, wouldn't it be faster to do it myself? But, no, it does end up being a lot quicker, but you aren't quite in that nirvana of fully, you know, autonomous acting agents that are, you know, roaming free amongst your corporate applications. And that's because of the verifiable ability and validity problem. When you can't assert upfront that you know exactly what this agent is going to do, then it becomes increasingly hard to, accept the risk of saying, I can let this go, and I know exactly what's gonna happen. And that leads into sort of the final problem that we introduced, which is determinism and fault tolerance. It's not true to say no organization can have completely free roaming, AI agents. It's a question of risk and fault tolerance. If your organization and the function that you're, deploying this agent for is a fault tolerant function, I. E. If something goes wrong, then it's recoverable. You can figure out what the problem was. You can remediate it. There's no kind of legal liability or outstanding risk. Then you're kind of in the right space to be using agents. And that's where, like, my job internally is I build products that use AI. I on a team of people that build products that use AI, and we hammer into all of our product teams the mantra that if you cannot tolerate persistent incurable error, don't use AI. And that's something that if you can't put up with it being wrong, AI is the wrong tool for you. So those are kind of the new problems, but there's some unique problems as well that we need to deal with. The first problem I'm gonna jump to is actually the kind of the last problem, which is AI can be expensive. And we're not just here talking about, monetary impact because there's kind of a race to the bottom at the moment in terms of token cost, which for those of you that haven't buried your head deep in the sand of AI world, is sort of the rating rating cost of how much you use an AI API. And we're kind of all saying, oh, look. Token costs are kind of falling year over year. But the truth is AI costs to companies is growing year over year. And that's because even though the cost of the last year's model is cheaper now, everyone's moving to this year's model and the cost of that stays about the same or even increases year on year as people release more and more capable and competent models. And I know Julian's gonna go into this in, in his section, but these AI agents start to chew up a lot of these tokens. They start to, do things where you would normally say, hey. I could have done that for for nothing. The AI agent just spent 15¢ listing out all of the directories in the folder so that it understands where it's at. You know? Cool. That doesn't sound like much. Right. But it did it 500 times today. So suddenly, I don't feel like I'm, you know, willing to pay $20 for an agent that's going to spend its time listing out folders. So they aren't always cheaper than people. They can be more expensive. They aren't always performant. And the last challenge that we see quite a lot of with AI agents is, they often require a human in the loop. So this there's this ongoing question of what's the true ROI that you get out of an agent. The return obviously improves as you can remove people from the loop and let the agent act more, agentically. But that incurs additional risk, which is, see, on the left hand side here that, you know, you now don't have anyone to correct it as it goes. So these are some of the unique problems that we see and some of the unique, challenges that occur with agents. Let's talk a little bit more about the next topic, which is the core underlying difference between the two. We're used to, like in Salesforce, in every application, the world of logic based. We all test our Apex apps. We've got Apex unit tests. I'm sure you do because Salesforce says you have to have Apex unit tests, and we would never find a way to gain that system. So you've all got your Apex unit tests. You've all got your functional scenarios. You've got everything set up. You know you, what's going in. You know what's coming out. We're all very comfortable in that world. And then they release agent force and say, cool. It works. Trust it. Let it go. Let it start building things. You have to make that jump between testing logic based systems and testing something that is more probabilistic. You need to go from the world of given this then that into the world of given this, probably that. And that's not a world that's, we're necessarily comfortable with if you didn't grow up in the world of data science. So let's move along to talking about, the correct places to, use AI and the correct ways to test AI and how this is going to link into how we use and how we test, Salesforce agents and agent force implementations. So the first thing to realize is that, your app can use AI and your app is AI, a kind of different things. So within any AI powered app, there's an AI and there's an app part. And you're either testing one or the other. You're testing the app, the logical part, or you're testing the AI, the fuzzy part. And the first confusion that people come up with is make making the assumption that these are the same, saying no. It's all one blob. It's an AI and it's an app, and it's not. And the best way to show that is, you think of an application kind of like this, and every Salesforce agent is exactly like this. There's the GenAI thing down the bottom, the the fuzzy monster, if you will, that's got all of its, fuzzy logic in it that does the analysis or the recommendation or the fuzzy matching or or looking at your app and determining what it should do. Like, all of that fuzzy logic is encapsulated in this nice blurry monster down the bottom here. And often you'll interact with it with questions and answers. So this will be the agent where you say, hey. Go do this, and it will pass that off to the fuzzy monster and say, what should I do? Then that answer will come back, and it'll be in the form of the agent telling Salesforce to do things. So you might say, someone's just come in with an email saying they wanna buy a product, but they don't have an account go, and the agent will say, cool. I'll go create an account for them. I'll create an opportunity. I'll convert to a lead. It'll do all of that work for you. But it's not going to do all of that work with the fuzzy logic. It's going to create actions and pass them back up to the the static logic part, where it's calling Salesforce APIs, and it's calling off to, you know, Apex implementations or whatever else. And the most important thing to remember is when you're testing these things, that you don't cross the line in the middle. So if this is, a testing strategy that you're developing for testing agent force in fact, this works for testing any agentic or AI based application, you don't put your functional tests on the fuzzy monster. Your functional tests, your Apex stuff, all the things that you're used to today still work and are still necessary and valuable, but don't apply them to the GenAI agent. And there's a few reasons for that. The most compelling, ones are that they just straight up won't work. The agent is going to give you different results for your functional tests. A simple one where you say, I'm gonna say hi to the agent, and I expect it to say hi back. Today, it will say hi back. Tomorrow, it will say hi. How are you? And the day after, it will say good to meet you because we expect a certain degree of that, you know, you know, unpredictability from the AI agent. So don't let your functional tests go over the line into your fuzzy monster. So that's something that's probably the most important thing. The second reason is because if you do let your functional test go over the line into the fuzzy monster, then you're gonna very rapidly incur a fairly heavy price because these functional tests are running an extremely powerful machine. In fact, we have some stats here about just how much it costs to do a simple prompt with a generative AI. I like to say that every time you send a hundred words to an early chat GPT model, it's like that, the battery in your phone oh, nice. My phone vanishes if I put it over there. But the battery in your phone, completely hundred to zero drained with those hundred words. That's how much power you're using. And in terms of the cost on the environment, they say it takes about three bottles of water. That's purely used in terms of cooling potential and processing. So it's expensive to talk to an AI. Now that may be valuable if it's doing productive work for you. But if you've just pointed your automated tests at your AI agent library, then your AI bill is going to go through the roof. So please adhere to the main principle. Do not cross the fuzzy line with your functional tests. So let's move on to talking about one of the reasons why you do want to test AI, because I've been, you know, fairly negative up till now saying, ah, testing AI is hard, you know, expensive, although there are a lot of problems with it. But there is a actual practical benefit to the rollout of your, AI agentic projects to having good testing in place. And that benefit is in terms of real concrete speed up to your, quality of release. Now when you're talking about AI, you're almost never talking about build costs. In fact, Ben said at the beginning of this, hey. The bill the testing cost was three x the build cost. You might say, well, that's that's insane. That's high. No. That's actually quite low. The median is it's about eight times. Testing and validation is about eight times the effort of building in an AI world. And the reason for that is, almost always, we are just palming all the hard work off onto the AI. We're saying, alright. Now here's a problem. Now go do the thing. Now the AI will come back to you on the first time and say, cool. The thing is done. And you'll go, great. I've solved the problem. Let's go to prod. Not so fast. You've got one example of it having performed that task. You've got a data point. You need to answer the question, how oftenly how is oftenly there you go. New word. How often, how well, how comprehensively does it complete that task? And if you don't have what's known as a benchmark, which is just AI fuzzy speak for a good set of tests, you don't know the answer to that question. You don't know what situations it's doing well, what situations it's doing badly. And this has a direct impact on your ability to deploy quality AI rapidly. Completion in an AI world is measured by how fast you can get to that benchmark level of quality. How quickly can you get to the point where you go confirmed. I know it's working to a degree where my customers are gonna get value out of it. Benchmarks are the key to unlocking not just the answer to that question, but speed in getting to that question. So here we've got actually an example. This is an internal project that we had at Tricentis where we were using a generative AI agent. And we didn't have, comprehensive benchmarking for about two months. It was about about two and a half months, actually. And during that time, we managed to get our product to a point where we felt like it was alright. It felt pretty good. Then we developed a benchmark. It took us a little while. We spent about two weeks developing the benchmark, so it wasn't cheap. And then we found out that it was actually only at about 46%, which married up with kind of the idea that we had of the heuristic that customers were giving us a feedback on. It took us one week from having that benchmark to be able to go from 46% to 93. Now remember that we had been working for two months to get to 46, and normally your, you know, your productivity should decrease over time. As soon as we got the benchmark, our productivity went through the roof. We gauge that we're about eight times faster after we had the benchmark at improving our quality than before we had the benchmark. So it is critically important to develop good tests for your agents, to have a good set of examples going in and expected results coming out so that you can answer that critical question, but also so that you can give feedback to the team that are deploying and prompting and building these agents to say, here's where it's doing well. Here's where it's doing poorly. And this shouldn't be considered as just a testing task. Although this is done by the people that know testing, this is a massive productivity task. Doing agentic work, doing AI work without a good benchmark dataset is the definition of insanity. So it is amazing. It is productive. It is necessary. It is difficult. That is the kind of summary about the your life when it comes to testing, generative AI agent based applications. So jumping over to the next slide, let's talk a little bit about one of the biggest challenges that comes with testing a generative AI, especially testing a generative AI agent. And that is to do not with the agent, the fuzzy logic itself. We've already talked about that to a great degree, what it is, the difficulty in testing a generative AI. But generative AIs are almost childlike in how they trust the tools that they are given. And when we talk about the tools here, generative AIs, remember, they don't actually open Salesforce and type things in the keyboard. They are like a brain disconnected from hands or eyes or anything else. So you need to connect them to the environment, and that's what's called tools. So you might connect a generative AI to a API for getting the weather, for example. And the generative AI, the agent, will say, hey. API, tell me what the weather is in this location. Now the problem is that the generative AIs have been trained to trust those tools. They don't question the answers they get from the tools. They don't question the quality of the tools. And so if that weather result comes back, like you say, hey. What's the current weather in Denver, Colorado? And it says back, oh, it's actually a 95 degrees, then the generative AI agent is gonna go, cool. Weather in Colorado is currently a 95 degrees, and it'll pass that back to the customer or the user. It's not gonna go, a 95 degrees. Everyone in Colorado is currently dead. It won't make that kind of adjustment. It doesn't question the tools and the results that it got back. And this has a pretty big impact when it comes to your Salesforce products. Because if you have defects in Salesforce, in the implementation of Salesforce, your normal users, your customers might pick those up by naturally questioning and saying, that doesn't look right. The AI agent is going to just trust it. It is going to take those results. It's gonna say, yep. That's fine. So if you don't test the tools that your generative AI agents are based on, then releasing generative AIs is going to create an explosion of defects within your product. So that's one of the major challenges with testing generative AI. It's also why you should be spending a lot of time actually testing the tools that the generative AI uses and not just the generative AI. The tools are the foundation. If you build on a shaky foundation, then it's not gonna be a fun time for the house that's above it. So the last thing that I wanna talk about here is how you benchmark ROI, AI. Because we've talked a lot about how you make sure that your agent is actually of good quality, but we haven't talked anything about do you even want that agent? Is it doing anything productive for you? And, ROI in agentic work is a little difficult to consider because you might and I'll I'll explain it with refer with reference to a product that we built, and then I'll link it into, into agent force for you. So we shipped a product that could create test scenarios for you. Great. Happy days. And it would get about half of those test scenarios were exactly the ones you wanted, half of them were ones that you didn't want. And we thought, great. It's, saving you about half your time. You're 50 more productive. The truth was that you were negative 25% more productive. And the reason for that was, sure, you save time on writing that first half of the test cases that you wanted, but you spent that time on reviewing the ones that you didn't want and determining whether or not they were correct or whether they were incorrect. And we're not used to the machine lying to us. So with generative AI with agents, you kinda have to make this trade off of determining what is the return that I expect, how much time am I going to be getting back, what is the deduction from that that I need to make with people going back and correcting offset here? And that gives you your return on investment. So you need to examine that from the point of view of how much you're spending today, how much, if you're using an AI agent, for example, to, interact with the chat people online and process orders for them or whatever you might be wanting to do, you might think, how much would I need to pay a human to do that work? And then that's kind of your ideal return because you're saying, hey. If I can offset 50% of that, then my ideal return is that I will take that 50% and subtract the cost that I'm paying for the AI, and there's my ROI. The only thing that I'm gonna outline for you here, because all of you know how to do that basic math, is you need to also consider within that cost the cost of repairing AI mistakes. So how often is the AI going to get it wrong and you need to delegate that over to a human to, repair? And how much more time does it take your huge human agent to go through and say, right. I've understood what the problem is. I can, you know, repair the damage that the AI has done, and also measure your, your customer success or failure with regard to the AI. Do they prefer it? Do they hate it? You know, how many sales do you win or lose? So thinking about ROI in, agent terms is much more of a business, decision. It may not always be the best idea to palm it all the way over to the agent. Sometimes a workflow or an automation might be a better decision for, getting the best yield out of your value. So the last thing that I wanted to, talk about is why you would bother testing agent force, because this is an out of the box product. Right? And so to talk about that, I'm going to hand over to the expert in the space, mister Julian Ho. Thanks, Dave. So, yeah, agent force unlocks a lot of value as we know, but it also raises lots of questions. So, as David sort of mentioned, how do you know it's producing the right results? If there's an issue with your agent, how do you even fix it? And, also, how is it interacting with your downstream Salesforce processes and data? And, also, about regressions. Like, why don't I introduce a a new set of problems, when I try to fix a fix a current problem? So to do this, you really need, like, a high quality agent. And as we've mentioned before, the best way to do that is to test that agent. And this is, I guess, the next main topic is really how do you test agent force. Commonly, we get asked, doesn't agent force testing center test agent force? So it does, but only parts of AgentForce, and it doesn't really test it in specific areas. So what we wanna do next is really explain to you how to completely test your AgentForce agents. So here's a an architecture diagram of Agent Force. So you have your agent, your topics that you've defined, and your actions, and then underneath that are your custom actions. So as David mentioned, these are basically the tools that underpin the agent. And then underneath that is your all your Salesforce processes and data. So, really, the software testing principles that apply to, typically to software also apply here in agents as well. And the principles are really around the testing pyramid strategy. So what we recommend is that you have, a a layered approach to your testing. At the bottom, you have many small, unit tests, very isolated tests that are very easy to write, very easy to run, and most importantly, very easy to detect issues and correct. And then as you move up those layers, you begin to have more integrated tests. They become actually harder to write, take longer to execute, and most importantly, longer to debug. So you have to think of it as a pyramid and specifically, as we've mentioned, around the tools underneath that underpin the performance of your agent. The other principle we recommend is around automation. So you should try to automate as much as possible. In that way, you get repeatable tests, and the tests also can be run more frequently. So you can always detect those issues and answer those questions that you may have about your agent. So as we mentioned, the tools, your custom actions are fundamental to the performance of your agent. So you really need to test those with very small tests, lots of them to ensure that the foundations are solid for your agent. Now in terms of testing your custom actions, it's pretty much like a patchwork because we are in the very early stages of some of these tools. So for flows, there is some basic debugging capabilities in Flow Builder. We know that more is coming later from Salesforce, so you have to spend a lot of time on this. But, there are capabilities, although quite basic around testing flows. We mentioned Apex and unit tests that everyone's aware of, which, of course, they have. And, look, they can be very easily added into a CICD pipeline to get that automation going. So, you're probably already doing the Apex tests. And in terms of prompts, again, similar to Flow, there's some basic debugging capabilities. So you have to sort of manually go in there and and test your prompts as well like flows. But we we really wanna emphasize that you need to test your custom actions like your tools that the AI agent is reliant on, and you need to make sure they're constantly working and they're they're of high quality. Then if we move up the stack, then we have the the agent actions. So they have to be defined to make sure that they're being called correctly and with the right inputs and outputs. In terms of testing agent actions, really the best way to do it or the only way, sorry, is really, in agent builder. By using targeted prompts, you're able to see that the correct action is being called and its inputs and outputs are correct. And then by doing it in agent builder, you you have to then, I guess, not guess, but just redefine refine your agent action descriptions. And by doing that, you need to iteratively keep on testing them to validate its inputs and outputs. And then if you want to test the whole agent as you move further up the stack, you need to have comprehensive prompt testing to ensure the agent's selecting the right topic and then those actions, and getting those expected results. And the way to do that is then to use agent force testing center to test the selection of the correct topics and the correct actions using prompts and then, again, refining the topic description, the scope, and instructions that you've given to the agent. Now Agent Force testing center, you can use GenAI to generate tests. You can use GenAI to actually validate your test results. However, what we recommend is that's more for the exploratory type testing you may wanna do, identify those corner cases. But in terms of your main testing, we recommend that you have a test spec file, either Excel or YAML. So you have very repeatable, very, specific testing that you can continually do. So the other point is also that you shouldn't rely on GenAI in terms of responses. So we talked about some of those issues at the beginning. So we recommend that don't rely on GenAI to validate your responses. You should probably use manual validation to validate those responses. And then the last bit is you wanna do an end to end test. So this is really at the top of the pyramid. You don't need to do many of these, but you want to validate that the agent is doing the correct actions on your Salesforce process and data. And the best way to do that is to use use a third party tool to do that. And that's where testing Salesforce comes in. It's our no code test automation product. It's low code. There's up there's prebuilt Salesforce steps to make it easy for anyone to write tests. Each test is metadata aware of Salesforce metadata, so these tests are very, stable and have low maintenance. It's underpinned by AI testing technology. And, also, we're always testing on the latest Salesforce release, so there's no need to have to maintain your tests for every new release from Salesforce. So what we'll do now is just do a quick demo of an end to end test using, testing Salesforce, which is our test automation tool. So here's the tool. I've said it's all no code, and what we're going to do is just run through an end to end test. So the first step is really just to pull up an opportunity that we wanna work on. So we're gonna ask the agent to update the amount and the probability of this opportunity and the budget confirmed, which will then trigger a a flow to generate a a new event. So using some random variables in in terms of our test data, we'll ask the agent to update this opportunity. We confirm that through our automation, and then we just refresh that. And then our testing automation tool will validate that they've been correctly done. So it'll check the amount and probability, and it'll check that that task was created. So that's a simple no code test, using testing Salesforce. You can continually run that test. I'll suppose that you can continually run this test. You can put it in terms of a a nightly nightly run. So you you continually testing the end to end quality of your agent from the agent all the way down to your processes and data within Salesforce. K. And just wrapping up sort of who who we are, we provide, enterprise testing solutions for for applications. We provide test automation for web, mobile, and Salesforce. We do performance testing. We have a a set of, test management tools and also test optimization tools around accelerating and reducing the time of execution for your tests. So altogether, we have an integrated engine, quality engineering platform. And to just wrap up about AI and Agent Force, as we said, it's very compelling and unlocks great value for customers and businesses. However, it's not free, and you have to really justify its costs. And importantly, what we want people to take away is need to ensure your agent and particularly your tools, your custom actions need to be working as expected. And you need to also consider downstream impacts, and that's the need to have a end to end testing strategy to ensure that the agent is working correctly with your whole Salesforce data and processes. Alright. I guess now we have some time for some q and a. Yeah. Cheers. That was super interesting. I think I learned a lot about not just testing, but kind of AI in general. So, yeah, thanks a lot for the insights there. Yeah. As Julia had said, we're gonna go into q and a now. So if you do have any burning questions, feel free to put them in the q and a chat box, and we will try and get around to as many as possible. I did see one come through from there, about is the, the x dollars per 1,000,000 token chart accurate? I didn't realize that there'd be a price difference by different LLMs. So, yeah, I can probably, jump on that, Gagan. So, the difference in cost between LLMs, is up to about a thousand fold difference in cost. So if you get a, like a a DeepSeeker one model, for example, you'll be talking maybe, like, 30¢ per million tokens. If you take an o three high, then you're talking, up to, like, $30, for the same thing. So you got huge differences there. Also, some of the more efficient, you know, Google models out there can be down around, like, the 1 to 10¢ per million tokens. So seeing these huge differences in terms of token costs is actually quite common, and it changes month to month. So, every time a new model is released, the typical pattern is that the old models will be discounted down quite significantly too. It's like, you're it's very common to see drops of, it going down to one tenth of the prior price, once the new model comes out, which is kinda why you see that massively wide variation. There is a a kind of ironic point where, GPT 3.5, for example, at the moment can be more expensive than GPT four o, because they just stop discounting the price after a while and forget about it. So everyone that's on those old models just kind of pays the price for their lack of attention. But that's, yeah, it is very common. It is a prevalent problem in the industry. And those people kind of think, oh, the cost of tokens is going down, so I can bet on this becoming cheaper over time. But they don't account for their own behavior that they constantly upgrade to the latest model because they go, oh, this is how I get the best value out of it. And so the cost of, ANSA remains relatively stable or in some cases goes up over time, especially with the new reasoning models, they tend to use about five times as many tokens to do the same task as a non reasoning model. So it's a bit of a a cost fallacy thinking our cost is just going to continually degrade over time. Yeah. Thanks a lot. There there is a great starting with API pricing, OpenAI API pricing into Google is a great, page where you can see all the pricing of the different models. So thanks a lot. And do you have any recommendations for ensuring that AI agents have minimal testability issues? So any approaches for apps with AI that make them a bit more testable? One of them that we covered in this session is always separate AI and not AI work. So if you've got an app that uses AI, usually, that AI can be boxed in. We call that the the fuzzy box. That's where your AI is working. That's anything that the AI agent is doing with a language model. Separate that from your tools so that you can make sure that the tools work. You can also separate it from logical flows. So if you've implemented a workflow, for example, you'll have conditions where if the AI says x, go x. If the AI says y, then go y. So you need to make sure that all of those are tested independently. We do have, there's actually a great webinar on our website from a guy called Chris Colosimo about virtualizing AIs, which is sometimes you can't pull the box apart and say, hey. You know, I wanna pull the AI out and just test the other stuff. So you can just, sit in between the AI API and the app and simulate the AI. So you you make the AI predictable by saying, when you ask this question, I'm just giving feedback this response all the time. So there's a bunch of ways of stabilizing that. The but once you've kind of gotten into that realm of, now I do need to test the agent, I need to test the the fuzzy part, the most important thing is to break it apart. So try to test the smaller areas. Like, if you've got an agent that's gonna perform multiple tasks, try to test each task individually. And make sure that when you're doing this, you're kind of splitting apart your, your tests by the functions that they're testing so that you can monitor where it's better or worse. So if you've deployed an agent, his primary job is to field customer questions, then break apart that test by the different functional areas or your different product lines or your different business types of questions so that you can examine where if there's a particular area of failure, you might be able to isolate it by saying, oh, it seems like on product x, it gives poor answers. So so let's dig into that. Without that kind of partitioning up of your test set, you're just gonna kinda be stuck with an overall number. And then, there's a there's kind of a joke amongst our AI engineering teams that there is no such thing as prompt engineering. There's just prompt sorcery because you're going to, take a prompt and you're going to summon a magic spell of different types of, you know, incantations you can give that hopefully make the AI behave better. Now because the AI companies don't like the concept of saying, well, this is all very unpredictable and fuzzy, they call this prompting guidelines. But, effectively, that is just, the incantation that you summon to try to make the AI behave properly. But if all you've got is one big number that you're trying to move, then you're just really throwing things at the wall and seeing what sticks when it comes to moving that number. If you can break it down into a small area, then you can target your prompt source or you just say, hey. Maybe I need to add something there to say, specifically, when you see this product, then go look in these locations to get it and teach the AI to behave a little bit better. Yeah. That's great. Makes a lot of sense. Thanks, David. And maybe on that note, how how would you approach using the Tricentis product sales source differently for automation with AI and AI agent flows? So I'll cover off the kind of best use case for the product and why it's crucial, and I'll pass over to Julian who can answer sort of the AI agent and AI agent flows part of it. The one thing the product does exceptionally well is it makes sure that those tools that your AI is based on are operating exactly the way you intend them to. So that is, probably the most important part about testing an AI is making sure that the foundations of the AI that it's built on are solid. When it comes to testing the actual agents that are sitting on top of that, I think Julian outlined that, pretty well in his section. So, Julian, do you wanna, jump into that? Yeah. So, yeah, you you can automate the testing of the agents. And I think what we just showed there is just how you do one end to end, test case. What what we suggest is that you actually do multiple and do them frequently and and have a real testing solution around, continuous testing. So as we've mentioned before, the agents are always changing. The data that they're using is always changing. So you need to continually test your agents. So frequent testing, as we mentioned, perhaps a nightly test to make sure that it's actually behaving as you intended it to do. So that's where, solar foundation is important, but also to have a continual testing, in terms of your agent. Great. Thank you very much. And just just a reminder to everyone who have a few questions, this will be recorded, and sent out afterwards. So if you wanna rewatch or send to colleagues, then, yeah, you'll be able to do that quite easily. And do do you have any recommendations of how you keep track of all the Salesforce processes and agent force can actually touch and ensure that you're hitting every single one of them when you're testing? Well, I guess that's that's really up to, it's up it's up to you to keep track of, like, you're bringing the tools, and as we said, you need to have all those unit tests around your tools. And I guess they're they're the flows and they're they're the things that the agent is touching. And that's why it's also important to have these end to end tests as well. So you need to have both, but particularly focusing on the tools and then also the end to end test to understand which which of those tools it's touching. So that's why we have that layered approach in terms of the pyramid. So understanding when your actions are being agent actions have been called, refining your topics as you go as you work your way up the pyramid. Great. Thanks, Julian. If I've gone just come in, I'm not sure I totally understand if Ariane, you might get guys. But how can we test in production, for example, after deployment to catch drift if the test mutates records? That's a that's a very good question, because when it comes to testing and production for a for generative AI, first off, we don't recommend testing and production for generative AI. You're not going to detect anything that you wouldn't detect in staging because you've already kind of surrendered to the madness of the fuzzy monster. So the way that you test in production is by monitoring. And when it comes to monitoring, you have to be very deliberate about the type of metric that you're monitoring for. So when you're talking about drift, the drift can occur in a couple of different ways, changes in customer behavior, changes in tooling, changes in models, all of the above. But the thing that you actually care about is the real productivity metric that the agent is responsible for. And this can be, a deliberate piece of feedback that was given by a customer saying, yes. This was great. It did the job. You know, happy days. There's gonna be a thumbs up, thumbs down. But those aren't really the best metrics because, first off, they're voluntary opt in metrics, and so you always get kind of sampling biases and everything else that goes along with them. But, secondly, they don't really say much about productivity. Someone could have gotten the job done and been like, yeah. I still, you know, I still had an issue. I'm still giving it a thumbs down. So the metrics that we recommend you track to determine your quality in production, which is really what testing in production is your quality in production, which is really what testing in production is about at its core, is you identify the final productivity metric. So how many like, if you are creating, accounts or leads, then you don't measure how many accounts or leads you've created. That's a that's a, a false metric. You would measure what your conversion rates are on those accounts or leads that are created by AI versus ones that are created by humans, for example. And that tells you the relative quality of the work that the AI is doing. So this is where you're kinda splitting the two things apart. You split apart the amount of work being done, which is how many leads are being created. Yeah. That's just a a unit of entropy to use the Newtonian physics sense. And the other one is the quality of the work. Now you kind of trade off between them. Like, a lot of low quality work isn't necessarily a good thing, but that's where you have to make the business decision. Like, how high do I want that quality bar to be set? And if I move the work bar and have the AI do more work, do I see a degradation in quality, or do they align linearly? And that's what you monitor in production is you find that trailing metric, which gives you the true idea of what the quality of the work that the AI did was, in terms of conversions, in terms of revenue, in terms of satisfaction of the customer retention, whatever it is that you actually value out of the work the AI did, and you measure that and you monitor that in production. And then if you start to see a drift, then that's when you go, okay. Well, something's changing here, and you go dive in and you take it into the testing environments. And that's critical because often your tests don't drift. You don't change your test scenarios that you're running your end to end scenarios. So your customer behavior, if you will, is always static in the testing environment. And if that customer behavior changes in production, then that's often a sign that you need to take that update back into the testing environment. Say, well, let's collect some of these examples we saw in production and bring them into our end to end test or bring them into our agent test to mimic that behavior in the testing environment so that we can then go back into that improvement and cycle of refining and tuning our AI agents. That's great. Thanks, Julian. Thanks, David. I can see one more question, but we do have just on the five minutes left if anyone else has the others. But, last one is quite broad. But what type of applications do you see as best uses for Genetec AI? Oh, that is quite a broad question. First off, it has to be an an application or a a use case where you can tolerate fault. That's kind of the the entry criteria into it. But with Intracentis, we actually prioritize use cases, not based on what people ask for, but based on where people spend their time. Because agentic AI is entirely a productivity function. It doesn't, like, an agentic AI isn't going to expose some new capability that you don't have today. It can only use the tools that you currently have at hand. What it replaces is human effort. It in fact, it's called, thought automation or cognitive automation, which is kind of the concrete term for it. But if you look at where people are spending their time, that might be where your customer's spending your time, where your staff's spending your time, you know, where your executive's spending their time, then those are the things that you target first. Because chances are if you prioritize your work by time spent, then your return on investment is naturally going to be higher. A bit of a false, dichotomy is saying, oh, you know, we've got this really cool thing that we could do with generative AI, And we asked a bunch of people, and they said that they really wanted it, so we're gonna go and do that. It's sort of the the shiny object, fallacy. So if you chase down that path, then you'll end up with a lot of things that people go, wow. That's cool, and then they never use because it's kind of like a a new and interesting edge case that they haven't thought of before. Often, it's the boring things that you do every day that generative AI proves to be the most productive and most valuable at. A best example of this is there's a ton of AI assisted coding editors out there now, things like, you know, Cursor and Codium and Codo and all of those. But if you asked a hundred devs and you said, what do you find the most irritating and boring part about your job? They're not gonna say coding. Right? That's, you know, the thing that they do every day. It's part of their job, part of their life. They actually enjoy doing that. And so focusing on what they do gives the most productivity. Productivity gives the most value, which gives the most return on your investment. So simple answer is pass the gate of fault tolerance, target things that you do every day. Great. Thanks a lot, David. I feel like that's a million dollar question at the moment, the best use cases because everyone's kind of wrestling with that. Yes. It is. Cool. Well, those are all the questions we have. I'll just monitor the chat to see if anyone else has any others. But otherwise, yeah, we can wrap up. So, Julia, David, thank you so much for joining us today. Super insightful. I learned a lot, and I've got a feeling this this isn't the last we're gonna hear about testing agents in Salesforce. So it's so great to kick off this conversation because, yeah, I'm not seeing much much content now, and it's obviously a super important topic. So, yeah. Any last words before we sign off, Julian? No. Thanks for having us, Ben. Yeah. My pleasure. David? Yeah. Thanks, everyone. If you, follow our blog on Tricentis, we're gonna be releasing and we are releasing a lot more, like, deep dive content on how testing AI works, how agents work, and that kind of thing just because one of our corporate values is giving back to the community. So, yep, if you're interested in the topic and you wanna find out more, there's a lot of stuff going up there. Fantastic. And one minute to start. Great. Thank you so much, guys, and thank you for attending today. See you at the next one. Bye, Guru. Thanks.