LIVE · SAT 8 AUG · 7–8 PM IST RETHINK AI ACADEMY · SESSION DOSSIER 001

What is a Forward Deployed Engineer, and how to become one.

One hour, no jargon. We'll follow one real business with one real problem, and watch an engineer fix it. Then I'll show you how to get this job.

Rashin PothanPresented by
7th Pillar InfotechCEO · Kochi
ReThink AIFounder & CTO · Arizona
THE BUSINESS AI belongs here IN PRODUCTION

I didn't read about this role. I do it.

Rashin Pothan
RASHINKOCHI
Emmanuel Onate
EMMANUELARIZONA
CEO, 7th Pillar Infotech (Kochi)  ·  Founder & CTO, ReThink AI (Arizona, US)
  • I'm Rashin. I've been running 7th Pillar Infotech since 2014, building software for clients in the US, Europe and Australia.
  • Over the last three years we've been building AI solutions, for our existing clients and for new ones.
  • Co-founder of ReThink AI in the US, with Emmanuel Onate — ten years in business and consulting, and now implementing AI across a variety of businesses internationally.
  • AI business process automation, and AI agents across sales and operations.
  • Voice agents that answer the calls after office hours, qualify the leads and book the job.
  • And BookingOps, our SaaS for the dispatch, follow-up, scheduling and booking behind it.
  • Mostly small and mid-sized businesses. And every system runs inside their own stack.
0yrs running 7th Pillar
0continents served
0AI companies in production
24/7voice agent live in the US

A Day in the Life of a Forward Deployed Engineer.

Easier to show you than to define it. So we're going to follow one — Arjun, working the whole job from eight thousand miles away — from the day he picks it up to the day it goes live. One business in Arizona. One problem. Six weeks. Watch what he builds, then watch what he says no to.

Case no.
2026 / AZ-04
Sector
Residential cleaning
Location
Chandler, Arizona
An illustration of Arjun at his desk, headphones on, mapping a process by hand
ArjunTHE ENGINEER · 8,000 MILES AWAY
EXAMPLE CASE BASED ON REAL PROJECTS
NAMES AND NUMBERS CHANGED

Fabulous Cleaners LLC

The cleaning company Mark started back in 2014, in Chandler, Arizona. Most customers book the same slot every week or every fortnight. They also do one-off deep cleans and end-of-tenancy cleans.

Who's in charge
MarkFOUNDER & OWNER
Started the company in 2014, cleaning houses personally, with one car and an ad online. Still knows half the customers by name, and runs the whole business from a phone.
Rosa Reyes38 · OPERATIONS
She sends 17 teams out every morning, sorts out who's called in sick, calms down upset customers, and answers most of the phone calls. If she takes a day off, the week falls apart.
Tyler24 · OFFICE ADMIN
Rings people back. Asks happy customers to leave a review.
The crews34 CLEANERS
17 teams of two, 12 vans. Same customers, same day of the week, every week.
WHAT THEY USE: Zoho CRM for customers · Google Calendar for the schedule · Google Sheets for pricing · a paper notebook on the desk · email · QuickBooks
None of it is connected to anything else.
FABULOUS CLEANERS EAST VALLEY · PHOENIX METRO
0earned in a year
0cleaners
0regular customers
0calls a month

It starts with a phone call at 10:52 on a Tuesday.

INCOMING · MOBILE
(480) 555-0147
RINGING… 4 RINGS
"Hi — yeah, I'm looking to get a quote? We're in Gilbert, near Val Vista and Baseline. Four bedroom, three bath, about twenty-eight hundred square feet. Two dogs. It hasn't been done properly in… a while. My in-laws land Thursday."
TV on. Dog barking. Toddler.
What actually happened next
Phone rings. Rosa is already on another call.
4 RINGS
It goes to voicemail
10:52 AM
The message sits in the inbox
—
Tyler gets round to the voicemails
+ 6h 00m
Tyler rings her back. She doesn't pick up.
4:52 PM
Tyler leaves a voicemail
1 MIN
Nobody at Fabulous Cleaners ever knew this call happened.
Result: nothing.
She booked with another company at 11:20 that same morning. They picked up on the second ring.
Here's the part that stings. Answering her would have taken ten minutes. But nobody was there to pick up, and the same thing happened 230 more times that month.

Everyone blamed the scheduling. But the money was already gone before scheduling.

Calls that came in, one month0
Somebody picked up0
↓ 230 calls nobody ever answered. That's 36%.
Were new people wanting a price0
Actually turned into a booking0
01
36% of calls went unanswered. Lunchtime, busy mornings, and every single evening and weekend. When your house needs cleaning you ring three companies and book whoever picks up.
02
About 1 in 5 callers speaks Spanish. They were told someone would call them back. Usually that meant tomorrow. Usually it meant never.
03
Rosa spent 3.5 hours a day on the phone while also running 17 teams. She was the safety net and the traffic jam at the same time.
04
9 booking mistakes a month. Two crews sent to one house, wrong address, no gate code. Each one wastes a van, two people and half a morning.
41% OFFICE SHUT
of the calls came in when the office was shut.
Nobody picked up a single one.
They were flat out busy and losing money in the same hour. Both were true.

He didn't start with code. He started by watching the business.

A Forward Deployed Engineer does two jobs at once. First they work out what the business actually needs. Then they build it. Same person, both jobs. That's the bit most people get wrong — and it is why Arjun spent nine days before he opened an editor.

STEP 1 · THE AUDIT
Watch how the work really happens.
Get on a screen share and watch them work. Listen to their old calls. Do the job yourself for a couple of days. Then decide what is worth building.
4 DAYSNO CODE YET
STEP 2 · EVALS
Write the evals before you write the code.
An eval is a test for an AI system. Real examples from the past, each one marked right or wrong by the person who does the job today. You score every version against it.
5 DAYSSTILL NO CODE
STEP 3 · BUILD IT
Build one thing. Score it against the evals. Switch it on slowly.
Build, test, look at what broke, fix that, test again. Then turn it on in the place where it can do the least damage.
4 WEEKSBUILD + LAUNCH
Most engineers start at step 3. That's why most AI projects quietly die.

Before writing a single line, he watched them work.

Four days · all of it remote
Two days working their enquiry inbox himself, from his own desk, 8,000 miles away
A whole Monday morning on a screen share with Rosa, watching her real screen
120 recorded calls, start to finish. Not skimmed.
Rang six customers who had left. The ones who quit tell you the truth.
Timed it. Counted it. What people think is happening is not data.
Everyone said the problem was scheduling. The recordings said it was gone an hour earlier.
What he handed over: a map of all 16 steps, with a decision written on each one
01
Call arrives
ALREADY THERE
02
Answer and talk
NEEDS AI →
03
Understand what they want
NEEDS AI →
04
Look up the caller
PLAIN CODE →
05
Work out price and slots
PLAIN CODE →
06
Put it in the calendar
PLAIN CODE →
07
Pick the team
LEAVE IT
08
Plan the day's route
LEAVE IT
09
Load the van
BY HAND
10
Team checks in
SETUP, NOT CODE
11
Clean the house
BY HAND
12
Quality check
BY HAND
13
Cover a sick day
A RULE, NOT CODE →
14
Send the bill
ALREADY THERE
15
Take payment
ALREADY THERE
16
Ask for a review
PLAIN CODE →
0nothing new
gets built
0just plain code,
no AI
0genuinely
need AI
Three of those steps are people scrubbing floors. So of the 13 steps that could have been software, he built six — and refused seven.
DETERMINISTIC · 4 OF THE 16 STEPS

Plain code means you can write the rule down.

No model anywhere near it. Same question in, same answer out, every single time. It is ordinary code, and you could sit down and read the whole of it end to end.

One real call, all the way through steps 04 → 05 → 06
04Caller ID, looked up in ZohoNOT A CUSTOMER
051,850 sq ft — the base rate$240
3 bed · 2 bath+ $55
every fortnight, not one-off× 1.00
inside the fridge+ $40
THE QUOTE$335
Is that ZIP inside the service area?YES
06First free two-cleaner slot → write it in → confirmation textTUE 9:00
Where the rules came from

He invented none of it. Every number was already in their pricing Google Sheet and the notebook on the desk. He copied them out. That is the whole trick — the business already knows its own rules, nobody had ever written them down as code.

Why not just let the AI do it

Let a model price a job and it starts being helpful. It rounds the square footage down. It hands out a discount nobody agreed to. And you will not spot it for months.

And how do you know it's right? You unit-test it.

A unit test is two lines you write once, that run again every single time anybody touches the code. It pins the rule above to an answer the business already agreed. It passes or it fails — there is no score, and no opinion.

1,850 sq ft · 3 bed · 2 bath · fortnightly · fridge
MUST COME BACK $335
Someone edits the pricing next March and it returns $340? Red on your screen in one second — long before a customer is quoted the wrong number.
Plain code gets unit tests. The two steps that need AI can't be tested this way at all — those get evals, and that's coming up.
NEEDS A MODEL · 2 OF THE 16 STEPS

These are the two you cannot write a rule for.

They fail as code for two different reasons. 02 is a hearing problem. 03 is a meaning problem.

02 Answer and talk

Nobody is reading from a script. Write the rule that survives all four of these:

"…hang on — DAISY! QUIET! — sorry, what did you ask me?"
"we're the one just behind the Circle K, uh…"
two people arguing about which day suitsboth on the phone
"buenas, quería preguntar…"switches, mid-call
03 Understand what they want

Four callers. Every one of them is asking for the exact same thing — a deep clean:

"it hasn't been done properly in a while"
"we're moving out Friday, I need it spotless"
"my mother-in-law is visiting"
"the usual, but do the oven this time"

Neither list is finite. Every if statement you write here just makes the next caller the one it doesn't cover.

■ WHAT YOU GIVE UP

It is probabilistic. The same call can come out differently twice. You cannot unit-test it, which is exactly why most teams ship it blind and find out from customers.

● WHAT YOU DO INSTEAD

You build the answer key first and score against it — that is an eval. And you decide up front where it has to stop talking and fetch a person. A complaint never reaches it.

STEP 13 · COVER A SICK DAY

A real problem. The wrong shape of fix.

The ask

"People calling in sick ruins our whole day. Can you build us something?"

The test

What would the software actually change?

Rosa already knows within minutes — the cleaner texts her before seven in the morning. She is not short of information. She is short of a second cleaner. An app would have told her faster about a gap it could not fill.

The fix

A standby list, and $40 to anyone who covers at short notice. Working inside a week. No code.

If the honest answer is "we'd know sooner" — and knowing sooner doesn't fix it — then it was never a software problem.

Every business runs a dozen processes. Two of them were worth AI.

The processTimes a monthWhat it costs them todayVerdict
Taking enquiries
640
36% never answered — 230 of them. Each miss books elsewhere.
AI
Quoting a job
410
Only the 410 that got answered. Hours of work, then hours of waiting that turn into days. First quote usually wins.
AI
Invoicing & payment
1,100
Biggest number here. QuickBooks already does it.
LEAVE IT
Working out the price
185
A formula, not a judgement call
PLAIN CODE
Slots & booking
185
Rules all the way down
PLAIN CODE
Crews & routes
26
Same crews, same days, 22 miles. Already fine.
LEAVE IT
Complaints & damage
11
Somebody has to own the apology
HUMAN
ENQUIRIES
= VOLUME
640 a month, a third on the floor. Fix the most frequent leak first.
QUOTES
= SPEED
First quote in usually wins. Speed is the conversion rate.
AI earns its place when a process is high volume, still done by hand, and the delay costs you the sale.

If you can write the rule yourself, don't use AI.

◀ PLAIN CODE — YOU WRITE THE RULES
Price = size × bedrooms × how often
Extra charges for the fridge, oven, windows
Is this address inside our area?
Which slots are free right now, from the calendar
Put the job in the calendar
Send the confirmation text
THE LINE
AI — YOU CAN'T WRITE THE RULES ▶
Understand someone talking over a barking dog
Work out that "it hasn't been done in a while" means deep clean
Do the whole call in Spanish
Notice this is a complaint, not a booking
Ask the one question that's actually missing

Let AI do the pricing and it starts handing out discounts nobody agreed to. It rounds the house size down to be helpful, and you won't spot it for months. The proper word for the left column is deterministic: same question in, same answer out, every single time.

Now try the other way round. Hand an upset customer to a set of if statements and you've earned yourself a one star review. Getting this line in the right place is most of the job.

Of the four things they asked for, he said no to three.

That isn't him being awkward. That is the job they were paying for.

"We want one system that replaces all of this."
The scatter is real. But a rebuild is a six-figure project that still would not answer one extra call. Fix the front door first, then decide.
NO
"We need an app so cleaners can check in."
Zoho already does this. Nobody had ever switched it on. Two hours of setup and training sorted it. Cost: nothing.
NO
"Can AI plan better routes for us?"
17 teams, same houses, same days, all inside a 22 mile circle. The routes are already about as good as they get. The software would cost more than it saved.
NO
"People calling in sick ruins our whole day."
A real problem, but software is the wrong shape of fix. They made a standby list and paid $40 to anyone who covered at short notice. Sorted in a week.
A RULE,
NOT CODE
The most useful thing he produced in his first week was a list of things he refused to build.

To answer that one call, you need six facts. They live in five different places.

What you need to knowWhere it actually livesReadable
Size, bedrooms, bathrooms
The caller says it out loud
✗
Price per sq ft, and the add-ons
Google Sheets. Three versions of it.
~
Is this already our customer?
Zoho CRM
✓
Are we free on Thursday?
Google Calendar
✓
Gate code, dog in the garden
A paper notebook on the desk
✗
Did they already get a quote?
Somebody's email inbox
✗
Nobody designed this. It just accumulated, one tool at a time, over twelve years.
So what did he do about it
One price list
Two of the three sheets deleted. One file, one owner, one place to change a price.
COST: AN AFTERNOON
Gate codes into Zoho
The field already existed. Nobody had ever used it. The notebook got typed up once.
COST: TWO HOURS OF TYPING
One way in
The agent reads Zoho, the calendar and the sheet through a single interface. It doesn't care where any of it sleeps.
COST: A FEW DAYS
He didn't move the data. He made it reachable. That's a week's work, not a migration — and working that out is a big part of the job.

The AI answers the phone. The clever part is knowing when to hand over.

● THE AI HANDLES IT ON ITS OWN
  • A price for a normal house inside their area
  • Moving or cancelling a booking
  • "What time is my cleaner coming?"
  • Gate codes, key box codes, "the dog is in the garden"
  • The whole call in English or Spanish
■ HAND TO A HUMAN, RIGHT NOW
  • Any complaint, or anyone getting annoyed
  • Anything broken or damaged
  • Big jobs: builders' cleans, or houses over 3,500 sq ft
  • Offices, or an address outside their area
  • Anyone who asks to speak to a person
  • Anyone who has asked the same thing twice
◆ MARISOL CHECKS IN THE MORNING
  • She reads every booking the AI made before the vans go out
  • Every single day for the first two weeks
  • After that, just spot checks
  • Every call is recorded and written down, so anyone can go back and look
This was never about getting rid of Rosa. She was missing the eleven calls that really needed her, because she was buried under the other six hundred. Now she gets those eleven.

300 real calls, marked by the people who do the job.

300 real recorded calls. Rosa and Mark went through every one. The person who does the job decides what counts as right. Not the engineer, and definitely not the AI.

0to practise on
0kept hidden
0the nasty ones
The 20 nasty ones · where systems actually fail
A telly, a toddler and a dog, all at once
An existing customer moving their slot
A salesman cold calling to sell SEO
A furious customer whose crew never showed
Spanish and English in one sentence
Asking for a service they don't offer
Six things to score · all agreed with Mark before any code was written
01
Worked out who was calling
≥ 97%
02
Price within $15 of Rosa's
≥ 95%
03
Never two crews to one house
MUST NEVER HAPPEN
04
Complaint to a person in 20 seconds
MUST ALWAYS HAPPEN
05
Handed over on everything in the red list
≥ 95%
06
Never promised what they can't do
MUST NEVER HAPPEN
They agreed what "working" means before anyone wrote code.
SHORT FOR "EVALUATION"

An eval is a repeatable test for an AI system.

You take the 300 real calls behind this slide. The person who does the job writes down the right answer for each one. Then you run the agent on the same calls and compare, one field at a time.

One call, checked field by field
The checkRosa saidThe agent saidResult
Who is calling?
new customer
new customer
✓ MATCHES
The price
$335
$340
✓ $5 OUT · RULE SAYS $15
Was that slot free?
already taken
booked it anyway
✗ GATE BROKEN
Why some checks are a percentage and some are absolute
● SCORED · THE SOFT ONES
A percentage, against a bar you agreed first
Being wrong occasionally is allowed, because Rosa is wrong occasionally too. You're aiming to be as good as the human, not perfect.
■ GATES · CAN NEVER HAPPEN
Not scored. It happened or it didn't.
One breach in 300 and the whole run fails, even at 99%. Two crews at one house isn't a worse score. It's a wasted morning and an angry customer.
AND THEN YOU LOOP Run all 300→ Read only the failures→ Fix the one thing causing most of them→ Run all 300 again↻ Until it clears 95% and no gate is broken
Only when it passes on the 60 calls it has never seen does it get to speak to a real customer.

The first eval run scored 74%. It sent two crews to the same house, and argued with an angry customer.

60%70%80% 90%100% PASS MARK · 95% 74% 87% 93% 95.6% 96.8% RUN 1RUN 2RUN 3 RUN 4RUN 5 PASSED
RUN 1 · FIRST TRY
74%
clashes 3 · handovers 71%
Failed in private, on the eval set.
RUN 2
87%
clashes 2 · handovers 79%
Check the number against Zoho first. Is this already our customer?
RUN 3
93%
clashes 2 · handovers 88%
Pricing moved to plain code. It was rounding down to be nice.
RUN 4
95.6%
clashes 0 · handovers 96%
Check the calendar when the phone rings, not at 8am.
RUN 5 · SHIP IT
96.8%
clashes 0 · handovers 100%
Never-break list clean. Now it can meet a customer.

The evals didn't just say it was bad. They showed him where. Half the first-run failures were one thing: existing customers treated as new leads. One lookup, moved to the front. Thirteen points.

Run it. Look at what broke. Fix that one thing. Run it again.

On day one, it didn't go anywhere near the main phone line.

What he built · one job, start to finish
Call comes
in
AI picks
up
Look them up
in Zoho
Work out
what they need
Price it
Check free
slots
Book it in
the calendar
Send a
text
Save the
recording
FIRST
Nights and weekends only
10 DAYS
It answers between 6pm and 8am, and all weekend. Those calls were going to voicemail anyway, so there is nothing to lose.
THEN
Only when nobody else can
14 DAYS
In the daytime it picks up anything that rings more than four times. Rosa reads every booking it made the next morning.
NOW
It answers first
ONGOING
Anything it hands over, a person takes. Nobody was asked to trust it. They watched it work on their own calls for a month first.
WEEK 1Watch and count
WEEK 2Write the evals
WEEKS 3–4Build, test, fix, test
WEEKS 5–6Nights only, then live
Six weeks. One engineer.
Start where the current option is voicemail. You can't do worse than voicemail.

Same crews. Same vans. Same area. Someone finally answers the phone.

Measured over the first three monthsBEFOREAFTER
Calls someone picked up
64%
99%
How fast it gets answered
4 rings, then voicemail
2 rings, any time
Calls outside office hours
none
all of them
Jobs booked from calls
74 a month
109 a month
Spanish calls dealt with there and then
almost none
all of them
Booking mistakes
9 a month
1 a month
Time Rosa spends on the phone
3.5 hrs a day
40 min a day
0
a year in extra repeat business
48 more regular customers than the three months before · each worth about $3,700 a year
0
a month from the extra jobs
35 more jobs a month, at about $285 for a first clean
0
a month handed back to Rosa
She runs the teams and looks after customers now, instead of living on the phone.
It costs about $1.90 per call answered. A receptionist costs $3,400 a month and can't work nights, weekends or in Spanish. It paid for itself in the first month.
"I've stopped waking up at 5am to check the voicemail."
— MARK, FOUNDER

The code took nine days. Watching took four. The watching was the real work.

What he builtOne job

Phone call, price, booking. Not the whole company. One job.

What he refused to build7 of the 13 that
could have been software

Including three of the four things they walked in asking for.

What they thought they wantedA system, an app, routing

None of it would have earned them a single dollar.

Anyone can write code now. The AI does that bit.

The hard part is knowing what's worth building, and being able to prove it works.

That's what a Forward Deployed Engineer does. It's also why companies can't find them.

Now let me give it a proper name, and then show you how to get there.

You've just watched one work. Now the definition makes sense.

Someone who reads a business like a consultant and ships like an engineer. A model can't tell you what's worth automating, or what has to stay human. That decision is the job.
● WHAT THEY ACTUALLY DO
  • Watch the real work, decide where AI belongs
  • Write the code, inside the customer's mess
  • Wire it into the tools they already pay for
  • Prove it with evals before anyone trusts it
  • Stay accountable once it's live, on real money
× WHAT THEY'RE NOT
  • Not a salesperson with technical vocabulary
  • Not support, sitting behind a ticket queue
  • Not an architect drawing diagrams for others
  • Not a researcher training models
  • Not someone who hands over a demo and leaves

The name: the military sends people forward, not kept at base. Palantir coined it. The AI labs hire for it now.

"Forward" means inside their workflow, not their building. We've never set foot in a client's office.

Two kinds of judgement that rarely live in one person.

◀ BUSINESS JUDGEMENT
How the work actually flows, not how the manual says it does
What it costs to keep doing it by hand
Who quietly benefits from nothing changing
What happens to the business when it gets one wrong
Whether the staff will actually use it on a Monday
FDE
TECHNICAL JUDGEMENT ▶
What a model can and can't be trusted with
Which systems it has to reach into, and how
What shape the data is really in
How it fails, and how loudly it fails
What it takes to keep it running at 2am
Most people only have one of these columns. If you write software, you already have the right-hand one. The left-hand column is the part you're missing — and it is completely learnable.

How to become one, in 30 days.

Four weeks · four checkpoints · one real business, start to finish
WEEK 1 · DAYS 1–7
Audit first. Then one job, start to finish.
◀ BUSINESS→
Watch the work, time it, price it, and decide what deserves to exist.
TECHNICAL ▶→
What a model and an agent really are, tools, and an unfamiliar API.
WEEK 2 · DAYS 8–14
Gather the golden data. Write the evals.
◀ BUSINESS→
Sit with the people who do the job. Agree the bar before anything is built.
TECHNICAL ▶→
Sample it, label it, hold a slice back, and decide how it is scored.
WEEK 3 · DAYS 15–21
Build the agents. Score them against the evals.
◀ BUSINESS→
Cost per run against what it replaces — and will anyone use it on a Monday.
TECHNICAL ▶→
Structured output, safe retries, context, and one knob at a time.
FINAL WEEK · DAYS 22–30
Defend it like an FDE
◀ BUSINESS→
The business case in the owner's numbers, and the CEO version of the story.
TECHNICAL ▶→
Build over what they already run, and the engineer's version of the story.
DAY 7
An audit, and a working agent
DAY 14
A golden set, and an agreed bar
DAY 21
An agent that clears the bar
DAY 30
A complete case study
What you walk away with · five artifacts a stranger can read without you in the room
ARTIFACT 01
The audit
How the work runs today, what it costs, and which processes you'd rebuild around AI. Including the ones you refused — that is the half that shows judgement.
ARTIFACT 02
The architecture
The stack, the tools, the guardrails — and a reason each one exists. Gates, logging, and how you roll it back.
ARTIFACT 03
The agents you built
The system itself, running on that business's real tools. Structured output, retries that can't double-charge, and it resumes after a failure.
ARTIFACT 04
The eval report
Real cases, the pass rate, what kind of failure each one was, and the rule for when it hands over to a person.
ARTIFACT 05
The business case
Hours returned, money protected, cost per run. Written in numbers the owner recognises, not in tokens.
On day 30 you shouldn't just understand the job. You should have evidence you can do it.
WEEK 1 · DAYS 1–7 · BUSINESS

Watch how the work really happens.

What you need to learn
  • Watching the work, not the manual. Screen shares, their old calls, and doing the job yourself for a couple of days. There is an SOP written down somewhere, and the real work differs from it. That difference is where the money is.
  • Mapping and timing every step. How often it happens a month, how long it takes, and who touches it. You cannot price a process you have not counted.
  • What it costs them today. Put a number next to each step, even an estimated one. A process nobody has priced always feels cheaper than it is.
  • A verdict on every step. Needs AI, plain code, or leave it alone — with the reason next to it. Most steps are not AI problems.
  • What you refuse, and why. Anybody can list what they built. The refusals are the half that shows judgement.
  • Who quietly benefits from nothing changing. There is always somebody, and they will decide whether this ever gets used.
● WHAT YOU SHOULD HAVE

A process map with a verdict on every step, including what you refused, and a number next to what the work costs them today.

DAY 7An audit, and a working agent
WEEK 1 · DAYS 1–7 · TECHNICAL

What a model and an agent actually are.

What you need to learn
  • What a model actually is. It predicts the next token — that is the whole mechanism. The context window, temperature, and why the same question can come back two different ways. Enough that it stops surprising you.
  • What an agent actually is. A model, a set of tools, and a loop. It decides, calls a tool, reads the result, decides again, until the job is done or it gives up. That loop is the entire idea. Everything else is plumbing around it.
  • One agent framework, properly. Pick one and stay in it. Tool calling, system prompts, message history, how a run is structured.
  • Giving a model tools. Tool schemas, how the model decides to call one, and what you do with the result it hands back.
  • Reading an API you've never seen. Auth, pagination, rate limits, sandbox keys versus live ones. This is most of the build in week one.
● WHAT YOU SHOULD HAVE

One job running start to finish — one real input, one real output, wired into their real tools, with nobody watching it.

DAY 7An audit, and a working agent
WEEK 2 · DAYS 8–14 · BUSINESS

Only they can tell you what counts as right.

What you need to learn
  • Getting the data out of the business. Recordings, transcripts, past jobs, whatever they keep. Somebody has to ask for it, and that somebody is you. This is the week the case study stops being a story and becomes data.
  • The person who does the job decides. You sit with them while they write down the right answer. Not the engineer, and definitely not the model.
  • Agreeing the bar with the owner. What percentage is good enough, in writing, before anything is built. Being wrong occasionally is allowed — the human is wrong occasionally too.
  • The things that can never happen. Which failures cost a wasted morning and an angry customer, and which are survivable. Only the business can rank those.
  • Where the handover sits. The written rule for when it stops and fetches a person. Agreed with them, not decided by you at midnight.
● WHAT YOU SHOULD HAVE

Real cases marked up by the people who do the work, and a written bar — the percentages you must clear and the things that can never happen. Agreed before you build.

DAY 14A golden set, and an agreed bar
WEEK 2 · DAYS 8–14 · TECHNICAL

Sample it, label it, and hold a slice back.

What you need to learn
  • Sampling honestly. Enough ordinary cases to be representative, and the awkward ones on purpose. If you only collect the easy ones you will pass your own exam and fail on the phone.
  • Holding a set back. A slice you never look at while building. It is the only number at the end that has not been quietly fitted to.
  • Scoring. Exact match, tolerance bands, and LLM-as-judge — including where a judge quietly lies to you.
  • Hard gates versus soft targets. The handful of things that can never happen, scored separately from the percentage. One breach in three hundred and the whole run fails, even at 99%.
  • Turning a rule into a threshold. "Hand it over when unsure" is not runnable. Confidence, cut-offs, and what the system actually checks.
● WHAT YOU SHOULD HAVE

An eval set you can run on demand, a held-back slice, and a scoring script that gives you one number and a list of failure categories.

DAY 14A golden set, and an agreed bar
WEEK 3 · DAYS 15–21 · BUSINESS

Does it earn its keep, and will anyone use it?

What you need to learn
  • Cost per run, against what it replaces. Tokens, retries and tool calls, priced against the thing it takes over. In the owner's terms, not in tokens.
  • Whether it justifies its own existence. Some things work and still are not worth running. That is a legitimate answer, and it is better delivered by you than discovered by them.
  • Whether the staff will use it on a Monday. A completely different question from whether it works. If it makes their day harder, it will be quietly routed around.
  • Turning it on where it can do least damage. Which slice of real traffic goes first, what you watch while it runs, and what makes you switch it back off.
● WHAT YOU SHOULD HAVE

A cost per run written next to what it saves, and a rollout plan that starts somewhere survivable.

DAY 21An agent that clears the bar
WEEK 3 · DAYS 15–21 · TECHNICAL

Build it so it recovers, then score it.

What you need to learn
  • Structured output. JSON schemas and constrained decoding, and validating the shape before you act on it.
  • Idempotency and failure handling. Safe retries, timeouts, backoff and dead letters. A retry that charges someone twice is worse than a crash.
  • State, resumption and tracing. Persist the run and checkpoint each step so it resumes from the middle after a failure — and log every model and tool call so you can explain any run afterwards.
  • Context engineering. What actually goes into the window on every turn, what gets summarised, what gets thrown away. Most "the model is stupid" bugs are context bugs.
  • Loop control. Step limits and stopping conditions. An agent with no budget will happily spend yours.
  • Retrieval, and when it earns its place. Embeddings, chunking, vector search and reranking — and the honest question first: could a plain database query answer this?
  • Changing one thing at a time. Prompt, context, tools and model are four different knobs. Turn two at once and you will never know which moved the score. This is the week you stop guessing.
● WHAT YOU SHOULD HAVE

An agent that clears last week's bar on cases it has never seen, and that you can kill mid-run and restart without breaking anything.

DAY 21An agent that clears the bar
FINAL WEEK · DAYS 22–30 · BUSINESS

Tell it to someone who signs the cheque.

What you need to learn
  • The business case. Hours returned, money protected, cost per run — in the owner's numbers, not in tokens and latency.
  • The audit write-up. The process map, the verdict on each step, and the things you refused with the reason next to each one.
  • The CEO version of the story. What it returned. Five minutes, no architecture, no jargon. Rehearse it out loud.
  • What you would do next, and what you would not. The second phase, and why it is not phase one. Knowing where to stop reads as judgement, not as a lack of ambition.
● WHAT YOU SHOULD HAVE

A business case and an audit a stranger can read without you in the room, and a five-minute story you have said out loud.

DAY 30A complete case study
FINAL WEEK · DAYS 22–30 · TECHNICAL

Tell it to someone who will maintain it.

What you need to learn
  • Building over what they already run. Their CRM, their calendar, their phone system. You do not get a greenfield, and an FDE who needs one is not much use.
  • The architecture document. Stack, tools, guardrails, where the gates are, what gets logged, and how you roll it back.
  • The eval report. Pass rate, failure categories, the never-break list, and the escalation rule.
  • The engineer version of the story. How it recovers. What happens at two in the morning when a tool call times out and nobody is watching.
● WHAT YOU SHOULD HAVE

An architecture document and an eval report that survive someone else reading them, and a system running on the business's own stack.

DAY 30A complete case study

The webinar hands you the map. The cohort is where you walk it.

The written roadmap goes to everyone on this call either way. The cohort is for people who would rather build the thing than read about it — 30 days, one real workflow, from audit to a deployed and evaluated system. We hand you the case. A real business, its recordings, its data and its mess — so you are not stuck looking for a client before you can start.

ONLINE · 40 SEATS
₹20,000
STARTS 1 SEPT 2026 · TWO SESSIONS A WEEK · TUE & FRI · 7–9PM IST
The full cohort, live and online. All sessions, all five artifacts, recordings, and the community.
STUDIO · 30 SEATS
₹30,000
EVERYTHING IN ONLINE, PLUS A DESK IN KOCHI
Same cohort, plus a desk at our Kochi office one afternoon a week. Build next to people doing the same work.
academy.7thpillar.com/fde-online Ends with a case study
that proves you can do the work
Free · as promised academy.7thpillar.com/webinar/resources The full reading, watching
and course list from tonight
Questions. Ask me anything — the roadmap, the tools, the job market, or what you should build first.
SPACE or → next  ·  ← back  ·  F fullscreen