Testing AI agents for failure modes before they break in front of your customers

Hey everyone,

I run structured reliability checks on AI agents/automations — basically stress-testing them against the things that don’t show up in a normal demo: contradicting inputs, missing data, stale records, edge cases at thresholds, tool calls that silently don’t fire (saw a thread on here about an agent that “answered” but never sent the WhatsApp message — that exact pattern is what I look for).

I’m not a Make expert in the platform-mechanics sense — there are people on here way better at blueprint debugging than me. What I’m good at is finding the specific scenario that makes an agent behave wrong before a real customer finds it for you.

If you’ve got an agent that’s about to go live (or one that’s already live and you’re not 100% sure it’s solid), happy to take a look and tell you what I find. Free for now while I get a feel for the kinds of setups people are running here.