AI Automation
LLM as a Judge Automation
Task decomposition
- An inbox is polled for new unread email (checked every minute).
- The incoming email is used to assemble a reply-drafting prompt with instructions on tone, what to decline, and when to ask for a call instead of replying directly.
- A response is drafted.
- A separate judge step scores the draft against five criteria and returns a score, a rationale, and specific improvement instructions.
- If the score clears a fixed threshold, the reply is sent automatically, no human review step in between.
- If it doesn’t clear the threshold, the judge’s feedback is folded back into the prompt and the draft is rewritten, looping back to scoring again.
Objective of the workflow
- This workflow tests a fully autonomous reflection loop: a generator drafts a reply, a separate judge model scores it against a rubric, and only a passing draft gets sent, with no human in the loop at any point.
- It’s a deliberate contrast to workflows built with an approval step, and it surfaces a real limitation worth being upfront about: as exported, the refine-and-rejudge loop has no hard iteration cap, so a draft that never clears the threshold could keep looping indefinitely. Built for CIS 515 (AI and Data Analytics Strategy).
Architecture
