You don't need a 60-page AI strategy deck. You need a checklist.

We've published a comprehensive guide on how to evaluate your company's workflows and identify AI agent opportunities. That post walks through the full methodology — the thinking, the frameworks, the organizational change management. It's thorough.

This post is different. It's the companion piece you print out, bring to your next ops meeting, and start scoring workflows against. If you want to evaluate workflows for AI agents without drowning in theory, this is where you start.

Why a Checklist Beats a Vague AI Strategy

Here's what we see repeatedly with mid-market and enterprise clients: leadership agrees that "we should be doing something with AI," a task force gets formed, and six months later there's a slide deck but no deployed agent.

The problem isn't ambition. It's that vague strategy doesn't translate into action.

A checklist forces specificity. Instead of asking "where could AI help?" you're asking "does this workflow score high enough on five concrete dimensions to justify a pilot?" That's a question you can answer in a single meeting.

Checklists also create a shared language. When your COO, your compliance lead, and your IT director are all scoring the same workflow on the same rubric, you skip weeks of misalignment.

Regulated-data checkpoint: If any high-scoring workflow touches customer data, financial records, health information, or audit-sensitive decisions, score the deployment boundary next — not just the ROI. Start with private AI for regulated enterprises vs public APIs (and watch the short version). For how gated multi-agent systems enforce that control in ops, see Project Cortex.

The Five-Dimension Scoring Rubric

To evaluate workflows for AI agents effectively, score each candidate workflow from 1 to 5 across these five dimensions. A workflow needs to score 20 or higher (out of 25) to be a strong pilot candidate. Anything between 15–19 is worth investigating. Below 15, park it.

Dimension 1: Volume

How often does this workflow execute?

  • 5 — Runs hundreds of times per day (e.g., invoice processing, customer query triage)
  • 4 — Runs dozens of times per day
  • 3 — Runs multiple times per week
  • 2 — Runs weekly or biweekly
  • 1 — Runs monthly or less
High-volume workflows compound ROI quickly. An agent that saves 3 minutes per execution across 500 daily runs is saving over 40 hours a day.

Dimension 2: Rule-Based Logic

How much of the workflow follows clear, documented rules?

  • 5 — Almost entirely rule-based with clear if/then logic
  • 4 — Mostly rule-based with occasional judgment calls
  • 3 — Roughly 50/50 rules and judgment
  • 2 — Mostly judgment-dependent
  • 1 — Almost entirely subjective or creative
Agents thrive on rules. The more a process can be expressed as decision trees or policy documents, the faster you can deploy.

Dimension 3: Data Availability

Is the data the workflow needs already digital, structured, and accessible?

  • 5 — Fully digital, structured, API-accessible
  • 4 — Mostly digital with minor gaps
  • 3 — Mix of digital and manual/paper inputs
  • 2 — Mostly manual or trapped in legacy systems
  • 1 — Data is scattered, unstructured, or doesn't exist yet
This is where many AI agent projects stall. A brilliant agent design means nothing if it can't access the data it needs.

Dimension 4: Risk Profile

What's the consequence of an error in this workflow?

  • 5 — Low risk; errors are easily caught and reversed (e.g., internal report formatting)
  • 4 — Minor operational impact; simple rollback
  • 3 — Moderate impact; may affect a customer or require manual correction
  • 2 — Significant impact; regulatory, financial, or reputational exposure
  • 1 — Critical; errors could cause compliance violations, safety issues, or major financial loss
Note: A low score here doesn't mean "never automate." It means you need stronger governance, human-in-the-loop controls, and — very likely — private AI infrastructure.

Dimension 5: ROI Clarity

Can you quantify the return on investment?

  • 5 — Clear dollar value per execution; easy to measure (e.g., cost per processed claim)
  • 4 — Strong proxy metrics available (time saved, error rate reduction)
  • 3 — Directional ROI estimate possible
  • 2 — ROI is mostly qualitative ("it would be nice")
  • 1 — No clear way to measure impact
High-ROI AI agents aren't the ones that sound impressive in a boardroom. They're the ones where you can point to a number before and after.

Printable Scoring Rubric

Use this table in your next workflow review session. List your candidate workflows down the left column and score each one.

Workflow NameVolume (1-5)Rules (1-5)Data (1-5)Risk (1-5)ROI (1-5)TotalPilot?
Example: Invoice matching5544523✅ Yes
Example: Strategic vendor selection122229❌ No

Scoring guide:

  • 20–25: Strong pilot candidate. Move to 30-day pilot plan.
  • 15–19: Worth investigating. Identify what would need to change to raise the score.
  • Below 15: Not ready. Revisit in 6 months or after infrastructure improvements.

The Do-Not-Automate List

Not every workflow should get an AI agent. Some things need to stay human — at least for now. Don't automate workflows that:

  • Require empathy as the primary output. Crisis response, employee terminations, sensitive client conversations. An agent can support these. It shouldn't replace the human in them.
  • Have no documented process. If nobody can explain how it works today, an agent won't magically figure it out. Document first, automate second.
  • Change constantly with no pattern. Some workflows are genuinely ad hoc. That's fine. Not everything is a system.
  • Carry regulatory liability without clear audit trails. If you can't explain to a regulator why the agent made a decision, don't deploy it in that context — or ensure you have the governance infrastructure to support it.
  • Are politically sensitive internally. If automating a workflow will create organizational resistance that blocks broader AI adoption, sequence it later. Win trust first.

Your 30-Day Pilot Plan

You've scored your workflows. You've got a candidate that hit 20+. Here's how to run a focused pilot in 30 days.

Week 1: Scope and Baseline

  • Document the current workflow in detail (inputs, steps, decisions, outputs)
  • Measure current performance: time per execution, error rate, cost
  • Define success criteria for the pilot (be specific — "30% faster" not "better")
  • Identify the human-in-the-loop checkpoints
Week 2: Build and Configure

  • Select or configure the agent (this is where context engineering matters — the agent needs the right information at the right time)
  • Connect data sources
  • Set up logging and audit trails from day one
  • Run internal dry runs with synthetic or historical data
Week 3: Supervised Deployment

  • Deploy with a human reviewing every agent output
  • Track accuracy, speed, and edge cases daily
  • Adjust prompts, context windows, and decision logic based on real results
Week 4: Evaluate and Decide

  • Compare pilot metrics against your baseline
  • Document every failure mode and how it was handled
  • Make a go/no-go decision on expanded deployment
  • Build the business case for the next three workflows on your scored list

When Private AI and Governance Matter

If your AI agent opportunities checklist includes any workflow that touches customer data, financial records, health information, or regulated processes, you need to think about where your data goes.

Public AI APIs send your data to third-party servers. For an internal report summarizer, that might be acceptable. For a claims processing agent handling protected health information? It's not.

Private AI deployment keeps your models and data within your infrastructure. It's not a luxury — it's a requirement for regulated industries. And governance isn't something you bolt on after launch. It's something you design into the agent from the start, with audit logs, explainability, access controls, and human oversight baked in.

This is exactly the kind of architecture we build through Project Cortex — AI-powered decision engines designed for executive-level accountability.

Start With the Checklist, Not the Technology

The organizations that succeed with AI agents aren't the ones that buy the fanciest tools. They're the ones that do the boring, disciplined work of evaluating their workflows first.

Print the rubric. Score five workflows this week. Pick the highest scorer and run a 30-day pilot. That's it. That's the whole strategy.

And if you want help doing it right — especially in environments where security, compliance, and data privacy aren't optional — book a Llama Research Blueprint Assessment. We'll work with your team to score your workflows, identify high-ROI AI agent opportunities, and design a deployment plan that meets enterprise governance standards.

Private AI. Secure Data. Trusted Intelligence.


For the full methodology behind this checklist, read our complete guide: How to Evaluate Your Company's Workflows and Identify AI Agent Opportunities.


Blueprint assessment: llamaresearch.ai · hello@llamaresearch.ai