Testing before the major changes
Before going live with a large change - a new model, capability, or a change in how the agent works - we evaluate the agent on a benchmark of scenarios that mirror how the Dema users interact with it: asking deep analytical questions, creating apps and dashboards, maintaining longer conversations where the goal takes shape over several messages. We grade the responses using a set of criteria: did the agent actually solve the problem, did it use the correct data and produce accurate conclusions, did it follow the internal rules, and did it stay consistent across the conversation. New models are compared side by side on identical scenarios before we adopt them.Monitoring in production
Quality is also measured continuously in production, where we track the trend of the performance and investigate anything that falls short, ensuring the high quality of the platform.If an answer ever looks inaccurate, use the thumbs down on the message. The feedback is stored with the
conversation and helps the team find cases to review.

