For Ross. Walk the five scenarios below and mark every point where you or your team would go look something up before replying. Each of those points needs a deterministic hook so Pam sees what you see.
Pull up to 400 open tickets per brand from Gorgias, keep ones where the customer spoke last, no Pam note yet, not junk. Caps: 24 iHR, 20 ITAM.
By Gorgias tag first, keyword match second. Five buckets: WISMO, Order Issue, Return Issue, Payment, Q&A. Zero AI here.
Verified order from Shopify (status, items, total), tracking from ShipStation with carrier events, Loop return state, live stock per SKU on the order (up to 4), store credits from the last 120 days, the customer's last 5 prior tickets, the playbook section for the category, approved learned corrections, live guidelines, and the refund authority ceiling. All read-only scripts, no AI.
Pam (Sonnet) writes the draft from the customer message plus the last 8 turns plus all the injected facts plus her rules. She can also fire her own read-only lookups if the scripts missed something.
Did the ticket close, did an agent already reply. Then post as an internal note with the paw header.
Your pam-pass/pam-fail tags override everything, then text similarity at 0.75, then a stale-order check. Only unresolved cases get one cheap AI judge call.
Shopify order, ShipStation tracking with carrier events, delivered/in-transit narration.
The current tracking scan (did it move since 5:30 AM), carrier delay notices.
Gap: 11 of the 76 "rep had better info" misses in the last 14 days were live tracking events that happened after Pam drafted.
Ross: when you answer a WISMO, what exactly do you open, in what order?
Order facts, items, fulfillment status, prior tickets.
Warehouse/fulfillment portal state, whether a reship or edit was already done, notes from a colleague, promises made in a live chat or earlier thread.
Gap: 10 of 76 misses were a promise or context from a prior conversation Pam's draft predated.
Ross: where do those promises live that Pam cannot see (chat transcripts, another inbox, someone's head)?
Loop return state, store credits (120 days), playbook.
Whether the label was already issued, return received/disposition, whether an exchange item is in stock right now.
Gap: 8 of 76 were Loop state changes; 11 of 76 were live stock (in or out of stock at send time, this also hits Q&A).
Ross: which Loop screens do you check that the API summary might be missing?
Order financial status, store credit lookup, refund authority ceiling with daily cap.
The exact refund amount already issued (including fees or partials) in Shopify/Loop, chargeback status.
Gap: 7 of 76 were a refund or credit amount already computed somewhere Pam did not see.
Ross: where does the final refund math actually live when it includes fees or partial returns?
Live stock for SKUs already on the order, playbook.
Stock for items NOT on the order, product page details, restock timing.
Gap: stock lookups today only cover the order's own SKUs; questions about other products get no stock hook.
Ross: how often do you check stock for an item the customer is asking about but has not bought?
Of the last 14 days' misses where the rep had better info, 66% (50 of 76) were timing, not missing hooks: the order state, tracking, stock, or thread moved between Pam's 5:30 AM draft and the human's reply hours later. A hook cannot fix time travel.
Two fixes work together: (1) add the missing hooks above so Pam sees everything a human sees, and (2) move drafting to reply time so the facts are fresh when the agent opens the ticket (the pilot on the plan page, week of 8/3).
Only 3% of triaged misses were Pam actually reasoning wrong.
Your answers place the hooks; the golden set you are labeling decides the referee. Both feed the same goal, drafts you can send as-is.