With your permission, we use analytics cookies to understand what is useful and improve your experience. You can browse normally if you decline, and change your choice in the footer at any time.

Privacy policy
Skip to content

A NEW CHAPTER

Magical is now backed by Shiva.

Shiva, founded by Lucas Marques, former COO of Méliuz, invests in companies built with artificial intelligence. We share a vision: turning the potential of AI into real results for every client.

Shiva

Venture capital
Founded by Lucas Marques

BEHIND SHIVA

MonasheesEndeavor Catalyst

Leading funds that invest in Shiva.

Read in Brazil Journal

AI agents in large companies: what needs to be in place before the first pilot

What a large company needs to define before its first AI agent pilot: the workflow, data and access, systems integration, human review and success criteria.

Three paths of blocks converge into a blue workflow passing through a checkpoint

A first pilot should help the business decide where it can reduce costs, expand capacity or protect margins. The pressure to “use AI agents” has reached many boardrooms before anyone has a clear answer to a simple question: an agent to do what, with which data, under whose control? Without that answer, the pilot becomes a demo. It works in the presentation, but never connects to the right system, the accountable person or the criteria that would say whether it is worth continuing.

This article is a checklist for leaders deciding whether, and where, to put a first agent into use. It is not a comparison of tools or models. It is about what needs to be ready in the operation so that a pilot ends in a decision rather than an impression.

Agent, automation or system: the distinction that shapes the pilot

“Agent” has become a broad label. In practice, it helps to separate three kinds of solution, because each one needs different groundwork:

  • Automation runs a known sequence of steps with fixed rules. If an order arrives complete, it moves on to the next system. It is predictable and easy to audit.
  • A custom system organizes a process that has no suitable tool yet, with screens, user roles, calculation rules and dashboards.
  • An AI agent interprets a question or situation in natural language, looks up sources and chooses the next step within defined limits. It makes sense when inputs vary and part of the work is understanding what is being asked.

Many problems pitched as “a job for an agent” are solved by automation or by well-written rules. That is not a failed AI project; it is a good diagnosis. And when an agent really is the answer, it usually depends on those same rules and integrations to work.

Why agent pilots stall at the testing stage

The reasons tend to be less technical than they look:

  • the process has no owner, so nobody is accountable for saying whether the agent got it right;
  • data is spread across spreadsheets, systems and conversations, and the agent only sees part of it;
  • there is no initial measurement, because nobody measured how much time or rework the workflow takes today;
  • the limits of action were never defined, so nobody feels safe letting the agent do more than answer questions.

Each of these can be addressed before the pilot starts. The sections below walk through that groundwork.

1. Choose a workflow, not a department

“Using AI in sales” is not a scope. “Identifying orders with lower-than-expected margins, based on this month’s sales and costs” is. A good first workflow has a clear start and end, happens often enough to be observed during the pilot, and has an owner who feels the problem every day.

Four questions help compare candidates:

  • Impact: if this workflow becomes faster or more reliable, what changes for the business? Revenue, capacity, cost or risk?
  • Feasibility: does the information it needs exist, and can it be accessed?
  • Risk: what does a wrong answer cost? Starting where errors are visible and correctable limits exposure.
  • Owner: who will follow adoption and say whether the result is good enough?

A high-impact workflow without accessible data is usually a better second step than a first one.

2. Data and access: an agent only knows what it can look up

An agent answers from the sources it can reach. If discounts and returns live in a spreadsheet that was left out of the integration, the answer about margin will be incomplete, and may still look complete. Before the pilot, list:

  • which sources the workflow actually uses today, including informal ones;
  • who owns each source and how often it is updated;
  • which data is sensitive and who is allowed to see it;
  • what the agent should do when information is missing: flag it, ask for it or hand the case to a person.

At Rodomilk, that was the starting point: cost and margin data was scattered across spreadsheets and paper records. Magical connected an AI agent to the carrier’s operational and financial databases, and management began querying costs and margins across a fleet of 25 trucks in natural language. Those queries are only useful because the databases are connected; without that step, the agent would be answering about part of the operation.

3. Systems integration: looking up is not the same as acting

The key integration question is not “which systems does the agent talk to?” but “what is it allowed to do in each one?” There is a big difference between:

  • reading data from the ERP, CRM or TMS to answer a question;
  • preparing a record, proposal or quote for a person to review;
  • writing or changing information directly in the system.

For a first pilot, reading and preparing are usually enough to show value at lower risk. Compatibility with each system (access, permissions, data formats) should be checked before it enters the scope, not halfway through implementation.

Business rules belong here too. An agent that drafts a proposal needs pricing, tax and margin rules written down explicitly. At RBR Transportes, the solution was a custom TMS, not an agent: cubic weight, taxes, fees, commission and margin are calculated in the same workflow, and a quote that took about 30 minutes in spreadsheets is now generated in seconds. The example shows what has to exist before any agent: if the rules only live in the heads of the people who quote, no solution will apply them consistently.

4. Human review and governance: set the limits before you start

In a pilot, governance does not need to start as a committee or a long policy. It needs to start as concrete answers written into the scope:

  • which actions the agent may take on its own and which require approval;
  • when a case goes to a person: missing data, an amount above an approval threshold, a request outside the scope;
  • who reviews the agent’s answers during the pilot, and how often;
  • how exceptions and errors are logged so the agent can be adjusted;
  • where data is stored, which vendors are involved and who maintains the solution after the pilot.

Sensitive commercial decisions, such as discounts, payment terms or quality acceptance, should follow the approval levels the company already has. The agent can gather the information and prepare the decision; accountability stays with the people who already own it.

From source to action, within defined limits
  1. Authorized sources

    Available data and access permissions are verified.

  2. Agent prepares

    Consults sources and organizes an answer or proposal.

  3. Person reviews

    Decides on sensitive cases, exceptions and missing information.

  4. Permitted action

    Execution follows the agreed scope and approval levels.

Missing information or an out-of-scope request returns for clarification. Preparing does not mean having permission to write to a system.

5. Define success before you begin

A pilot without agreed criteria ends in opinions. Before starting, write down:

  • Initial measurement: how much time, rework or effort the workflow takes today, measured the same way it will be measured afterwards.
  • Primary measure: one or two measures tied to the problem that justified the pilot, such as response time, rework or volume handled by the same team.
  • Quality: how to check that answers are correct, using a sample reviewed by people who know the process.
  • Adoption: whether people turn to the agent in their daily work or go back to the old way.
  • Decision: which result justifies adjusting, expanding to another area or stopping.

Also keep what was observed separate from what is projected. An estimate of annual savings can help set priorities, but it only becomes a result once it is measured in the operation. Connect time freed up to capacity, cost per interaction, margin or revenue where measurable; time saved does not automatically mean cash savings.

A pilot ends in a decision
  1. Measure first

    Record the initial measurement and success criteria.

  2. Test one workflow

    Observe adoption within a bounded scope.

  3. Compare

    Evaluate quality, time, rework and adoption.

  4. Decide

    Adjust the pilot, expand or stop.

When adjusting, return to testing with clear criteria. Expansion depends on evidence and assessing the next workflow’s dependencies.

What to avoid in a first pilot

  • starting with the tool before choosing the workflow;
  • picking the company’s most critical process as the first test;
  • letting the agent write to systems before the quality of its answers has been validated;
  • measuring only how testers felt, with no initial measurement for time, errors or volume;
  • expanding to other areas before the first workflow has reached a clear decision;
  • treating a successful demo as proof that the operation is ready.

Six questions for the internal conversation

If your company is deciding where to start, these questions help structure the internal discussion before bringing in any vendor:

  1. Which workflow, with a clear start and end, do we want to improve?
  2. Who owns that workflow, and who will review the results?
  3. What data does the workflow use, and where does it live?
  4. What will the agent be allowed to read, prepare and write?
  5. In which situations does a person need to decide?
  6. What do we measure today, and what would need to change to make continuing worthwhile?

At Magical, the work follows the same logic: it starts with a conversation about the operation, selects the most actionable opportunity with the company, defines a pilot in one area and uses the evaluation to decide whether to adjust, expand or stop. The steps are described in how Magical implements AI, and the kind of agents we build is covered in AI agents for enterprises.

A conversation about your business

Where could your company
become more efficient or grow?

Talk to a partner about opportunities to reduce costs, protect margins and increase capacity.