Why do written policies and actual decisions drift apart?
A credit policy, a claims manual or a compliance procedure is written in prose for people. Each person reads it a little differently, edge cases get decided by whoever is on shift, and the document itself changes without anyone re-checking the cases already decided under the old version. The policy says one thing; the case files show another.
Automating the decision with a language model alone moves the problem rather than solving it. The model reads the same prose and produces a plausible answer, but two runs can differ, and nobody can point to the rule that fired.
Policy evaluation automation closes the gap by making each rule explicit and testable: written so a policy owner can read it, compiled so it runs the same way every time, versioned so a decision can be tied to the rule in force when it was made.
How do AI agents evaluate a case against business rules and decision policies?
On MightyBot, policies are written in plain English and compiled into deterministic checks. Agents assemble the facts a rule needs from documents and systems, with a pointer to the source of each fact, and evaluate every rule. The output lists each rule as passed, failed or missing evidence, with the values used.
Policy Profiles let one workflow carry variations by jurisdiction, product or counterparty, so a state-specific rule or a customer-specific threshold applies without a second workflow. Before a policy change goes live it can be backtested against past cases, so the owner sees which decisions would change.
Exceptions route to a person with the evaluation attached, and the record keeps who ran what and when, which is the question compliance asks first. Where a judgment call is needed, the platform can call a model for that step and record its output as one input among the deterministic checks.
What do regulators and standards expect of rule-based decision systems?
The April 2026 interagency model risk guidance draws a useful line. Its definition of "model" excludes "simple arithmetic calculations, such as those found within spreadsheets, as well as deterministic rule-based processes and software where there are no statistical, economic, or financial theories underpinning their design or use." Deterministic policy checks fall outside model validation; the controls that apply are the ones for policies, change management and records.
Those controls are well defined. The OCC's Internal Control handbook asks examiners to "Determine whether policies and procedures exist to ensure that decisions are made with appropriate approvals and authorizations." NIST's SP 800-53 control CM-3 requires an organization to "Review proposed configuration-controlled changes to the system and approve or disapprove such changes," and AU-3 requires audit records that establish "What type of event occurred," when, where, its source, its outcome, and the "Identity of any individuals, subjects, or objects/entities associated with the event."
For automated decisions that reach customers, the bar is rising. The EU AI Act requires that high-risk AI systems "shall technically allow for the automatic recording of events (logs) over the lifetime of the system," and NIST's AI Risk Management Framework asks that "Processes for human oversight are defined, assessed, and documented."