The full loop
Five hour-long interviews, one Zoom link all day. LP questions in STAR format everywhere, one 20-minute presentation (Deck tab), and possibly a whiteboard design session on Bluescape. You've passed everything they've put in front of you so far.
1 · Same Zoom all day — join at each start time, not early (waiting-room requests die after 10 min; if timed out, "Try again").
2 · Never repeat a story. The five interviewers debrief together and compare notes. Fifteen stories ÷ five interviews = fresh material everywhere. Mark stories "used" on the Sheet as the day goes.
Zoom — one link, all five
Join the interview →Camera on the whole day. Shadow interviewers may appear — a tenured interviewer still does the evaluating. Rejoining between sessions may show a stale date on the meeting — ignore it, same link is valid all day.
The five sessions — who, and what to bring
Mike George + Kirstin Biever
Principal SA (ex-CIO, security/AI-ML, Salt Lake City) + AE Nonprofits (ex-Director of enterprise CRM, Ohio State)
- George writes constantly on agentic AI for nonprofits: evals ("beyond vibes"), per-request cost, governance. Talk genAI in terms of measurement, cost, guardrails — never hype.
- Biever ran CRM inside a huge mission org — she's lived the customer's side. Plain-language business framing and empathy for small IT teams land here.
- SA + AE pairing smells like a customer-scenario interview: work backwards from the mission, translate tech to outcomes.
Lead stories: #8 estimator workflow (CO) · #1 validation speedup · #14 critiquing a metric (his "evals not vibes" language)
Break — eat real food
Away from the desk. Glance at the 13:00 session card when you're back, nothing more.
John Streit
Sr. Manager, US Consulting Partners — likely ex-KPMG robotics/ML lead + agency CTO (identity partly unverified — check his LinkedIn Wednesday)
- Partner-org manager → he'll probe scaling yourself through others: enabling partners/SIs instead of doing everything solo.
- Consulting pedigree → how you'd scope a delivery: requirements, tradeoffs, who does what between AWS, partner, customer.
Lead stories: #10 dept-wide materials (enablement at scale) · #7 solo PoC (scoping) · #11 union stand
James Bland + Rafia Tapia
WW NGDE Principal PSA (DevOps/SRE, ML doctorate — you've met him, he advanced you) + Sr. SA (29 yrs IT, databases → blockchain → DocumentDB → genAI; teaches "genAI for nonprofits")
- This is the technical deep-dive: Bland on the ops/scaling axis, Tapia on the data/AI axis. Expect the whiteboard here if anywhere.
- Tapia is a database specialist — defend every data-layer choice (why this store, why vector search, how it scales).
- Bland already heard your PS1 stories on Aug 21 — bring the new ones.
Lead stories: #1 5× speedup · #12 misalignment paper · #15 porting my own product to AWS (CDK, live right now)
Break — walk, water, reset
Two more. The 15:30 one is the lightest; the 16:30 one is likely the hiring manager.
Thomas Grimes
AE Nonprofits (thin public record — possibly ex-Salesforce/Verizon public-sector seller; unverified, don't assume)
- Second AE lens: Customer Obsession and Earn Trust — handling ambiguous asks, tight budgets, saying no well.
- Know the nonprofit money levers: AWS Imagine Grant, TechSoup credits — AEs live there.
Lead stories: #5 working backwards from sellers (CO) · #13 co-author disagreement · #9 full-stack in four months
Madhu Bussa
Sr. Manager, Nonprofits SA — almost certainly the hiring manager. Ex-Change Healthcare data-platform director, Vanderbilt MBA, rose IC→manager inside WWPS. Publishes on data monetization + genAI agents.
- Hiring-manager bar: Deliver Results, Ownership, self-sufficiency across many small accounts; his team scales through reusable content and workshops.
- Sweet spots: turning data into mission value, sustainable architecture for cash-poor orgs.
- He respects career-changers with business acumen — your pivot is an asset here, frame it that way.
Lead stories: #2 abstraction layer · #6 the failure · #4 backbone→commit · save one spare for "another example"
Story assignments are defaults, not law — answer the question actually asked with the best unused story. Cross-cutting note from the research: George, Tapia, and Bussa are all publicly invested in genAI-for-nonprofits — one crisp, measured point of view (cost, evals, governance) serves in four of five rooms.
Day-of timeline
- 09:30Normal morning. Optional 15-min voice warm-up (Voice tab #6).No new material today. Check WhatsApp for Sabri intel.
- 10:15Tech check: Zoom test, headset, camera, screen share the deck once.Deck open in its own window · this site on the Sheet · phone charged, ringer up.
- 10:50Read the Sheet aloud once. Water bottle. Bathroom. Door closed.
- 11:00Session 1 — George + Biever. Camera on, smile.
- 12:00Lunch — real food, away from the desk.
- 13:00Session 2 — Streit.
- 14:00Session 3 — Bland + Tapia. Technical deep-dive; Bluescape possible.
- 15:0030-min break — walk outside, no screens.
- 15:30Session 4 — Grimes.
- 16:30Session 5 — Bussa. Finish strong — hiring manager. Have your questions ready.
- 17:30Done. Walk it off, text Sabri, eat something great.Jennifer dos Santos (RBP) takes over after the loop for outcome/comp.
Before Thursday — confirm these
Your questions — tuned per room
- George/Biever: "Where are membership nonprofits actually adopting genAI vs. where is it still hype?" · "How do account teams and partner SAs split an engagement?"
- Streit: "What makes an SA great to work with from the partner org's side?"
- Bland/Tapia: "What's changed in how nonprofits buy AI since we spoke in August?" · "What data problems do you see most at membership orgs?"
- Grimes: "What separates the SAs your customers ask for by name?"
- Bussa: "What does success look like at 6 and 18 months?" · "How does the team scale itself across so many small accounts?" · "What's your bar for an SA's first workshop?"
If things go wrong
- Waiting room past 10 min → "Try again," then email Casey, phone in view.
- Zoom drops → rejoin same link; dial-in +1 833-548-0276 as backup. One sentence, move on.
- Mind blanks → glance at the Sheet. That's its job.
- A session runs long/short → their clock, their problem. Stay even; next session is a fresh jury.
- A session feels bad → it probably wasn't. Interviewers who push hardest often score highest. Reset at the break — five independent votes, not one running tally.
Story bank — 15 stories
Six from mtl.ai, three from Explorai, two from teaching, three from research, one from right now. Cues, not scripts. [Red text] = details only you know — fill them in the notes boxes (autosaves; your PS1 drafts are already loaded where they carry over). Tap used on the Sheet during the loop so you never repeat one.
📋 The Sheet — keep open all day
Numbers to say out loud: 5× · UEFA Euro 2024 + Champions League · 3 providers → 1 interface · 7 years, thousands of students · 2 ICML workshops + NeurIPS under review · assessment in ~1 week from zero · my own app: CDK Phase 0 shipped this week.
Openers
"Tell me about yourself" (60–90s, you'll give it up to 5 times — vary it, don't recite):
"Three chapters. Seven years teaching math at Dawson College — calculus and statistics to thousands of students — where I learned to explain hard things to any audience. Then industry: at mtl.ai I built real-time computer vision that ran live during UEFA Euro 2024 and the Champions League, and led an LLM integration layer unifying OpenAI, Google, and Anthropic behind one interface. After a full-stack AI role building document analysis for construction bidding, I'm now a grant-funded AI-safety researcher — our paper on emergent misalignment was accepted to two ICML workshops. And because actions beat words: I'm currently migrating my own production app to AWS with CDK. This role is my two halves in one job — deep technical building and teaching — for organizations whose mission I actually care about."
"Why AWS / why this role?" — builder + teacher in one job · genAI depth exactly where nonprofits need guidance (cost, evals, governance — not hype) · mission-driven work is a pattern for me, not a pitch.
▶1 · The 5× validation speedup — mtl.ai
S — Real-time CV models (detection + segmentation) for live soccer — virtual ad replacement that ran during UEFA Euro 2024 and the Champions League. Model validation had become the deployment bottleneck.
T — I took on finding out exactly why — ship cadence depended on it.
A — Profiled the pipeline end to end instead of guessing; identified the key bottlenecks — [name the top one or two precisely] — fixed them [what you changed], measured before/after on identical workloads.
R — 5× faster validation → materially faster deployment cycles. [Ship cadence before vs. after if you have it.]
Probes: how did you profile · biggest single win · how did you prove results didn't change · what would the next 5× take
▶2 · The LLM abstraction layer — mtl.ai
Bank questions this answers: Invent & Simplify #1 (complex problem, simple solution) · #3 (made something simpler for customers — the team) · #6 (novel decision, big impact) · Think Big #1 (saw a chance to do something bigger than the initial focus).
S — Production assistant app for marketplace sellers; the LLM market shifted monthly under it.
T — I led the integration layer. Goal: model choice = config, not an engineering project.
A — One interface over OpenAI/Google/Anthropic — normalized streaming, tool calls, errors; per-task routing with fallbacks; eval harness before any swap.
R — Swaps in hours; cost routed to cheapest capable model [real % or $ if you have it]; team shipped features provider-blind.
"It's the problem Bedrock solves as a managed service — I built the small version, so I know why it matters."
▶3 · The regional threshold controller — mtl.ai · ⚠ the deck story
This is your presentation (Deck tab). Whoever hosts the presentation gets it there — don't also spend it as an LP answer in another session unless asked directly. As an LP story:
S — Two weeks before UEFA Euro 2024, a new billboard ad appeared featuring a fake soccer ball. Our production detection model fired on it. In a live virtual-ad-replacement pipeline a false ball detection is a visible artifact on air, to millions, with no retake.
T — Kill the false positives before kickoff without degrading anything else. I was tasked with building the fix.
A — We weighed retraining with the new ad in the data: too slow for two weeks, and no guarantee performance stayed consistent everywhere else. Chose a regional, class-aware threshold controller instead: raise the ball threshold only in the parts of the frame where that ad shows, only when it shows. Reversible, contained, testable. Then: talked to the operators first, the people running the broadcast system at the stadium, learned their controls (an MPC-style pad surface) and their reality: about 2 minutes of lead time before an ad rotates in, so tuning had to be live and manual. Changed the model's output layer from one merged detection class to per-class output. Added a controller layer at the input of the broadcast system: thresholds by class and by stadium region. Operators pre-define regions about 1 hour before the match from a priori knowledge of where problem ads appear, then raise the threshold live when the ad rotates in.
R — Zero false positives across every Euro match. Performance everywhere else stayed at expectation. Shipped inside the two-week window.
Retro — [Your line. Candidates: expose per-class output from day one; automate region proposals from the ad schedule; feed operator tunings back as training signal for the next retrain.]
Bank questions this answers: Bias for Action #1 (calculated risk, speed critical) · #2 (tight deadline, couldn't consider all options) · #5 (move forward vs gather more info) · Deliver Results #1 (tight deadline, what sacrifices) · #2 (unanticipated obstacle) · Invent & Simplify #1 (complex problem, simple solution) · #5 (usual approach wouldn't address) · Customer Obsession #3 / #5 (operators' real workflow).
Probes: why not just raise the threshold globally · how did you validate the controller before air · what if an operator missed the 2-min window · what did retraining actually cost in days
▶4 · The architecture fight — mtl.ai · fill the specifics
The shape that scores: respectful challenge + data + the decision went somewhere + you committed genuinely. The commit half matters as much as the backbone half — spend a third of the answer there.
Your candidate — a technical/architecture disagreement at mtl.ai: [which decision? what did you argue, with what data? which way did it go — and what did committing look like in your actual behavior?]
If it resolved in your favor, also have the version where it didn't (story 13 can be that one). Voice prompt #2 builds this with you.
▶5 · Working backwards from sellers — mtl.ai · fill the specifics
The signal Amazon wants: you changed course because of what users actually needed, at some cost to your own plan.
Your candidate — building the assistant app for online marketplace sellers: [a moment seller feedback/behavior contradicted what you'd built or planned — what did you learn, what did you change, what did it cost you, how did sellers respond?]
Say the realization moment out loud ("we'd built X; watching sellers we realized Y"). That sentence IS the story.
▶6 · The failure — mtl.ai · must be real
Asked in nearly every loop — with five interviewers, possibly twice (have a second, thinner one ready). Scored on honesty and changed behavior, not the failure.
- Pick something real — a late project, a wrong technical bet, a model that underperformed in production. "I work too hard" is an anti-signal.
- Structure: what happened → my share, no blame-shifting → concrete lesson → what I do differently now, with an example of the new behavior working.
- 60 seconds. No wallowing, no spin.
▶7 · The solo floor-plan PoC — Explorai
S — Estimai generates construction bids from architectural plans. Finding structural elements — column icons — across multi-page plan sets was manual and error-prone.
T — Nobody owned it. I proposed a PoC and delivered it single-handedly.
A — The bet: don't rasterize and throw a vision model at pixels — the PDF already contains vector primitives. I extracted them and built detection + cross-page association directly on vector data: simpler, cheaper, more precise. Reversible if wrong — that's why I could move fast.
R — Working PoC [timeframe]; [what happened next — adopted? shaped direction?].
Probes: what did the vector bet trade away · accuracy measurement + failure modes · path to production
▶8 · The estimator's real workflow — Explorai · fill the specifics
Your second Customer Obsession story — keep one per AE session (this one for George/Biever, #5 for Grimes, or however the day flows).
S — Construction estimators build bids under deadline from hundreds of plan pages. We were building AI document analysis for them.
A — [How did you get inside their actual workflow — watched them work? sat with a real estimator? what surprised you (what they actually count, what they ignore, where errors hurt)? what did you build differently because of it?]
R — [The feature/decision that came out of the field truth, and how estimators responded.]
For Biever especially: this is her life — a mission org's under-resourced staff and the tech people who do or don't understand them.
▶9 · Full-stack in four months — Explorai
S — Jan 2026: I joined Explorai as an AI full-stack engineer — after years as an ML/R&D specialist. New stack, new codebase, a shipping SaaS product with real customers.
T — Contribute across the whole stack, fast — not just the ML corner I knew.
A — [How you ramped: what was new (frontend? infra? their framework choices)? your system for learning it — reading the codebase, pairing, shipping small first?]
R — Contributed across the full stack to a live product within 4 months, plus the solo PoC (#7) on top.
▶10 · Materials adopted department-wide — Dawson College
S — Seven years teaching calculus, linear algebra, stats, probability to thousands of students. [The driving problem: where were students failing with the existing materials?]
T/A — Rebuilt the materials around how students actually got stuck — [what you changed structurally, and how you knew it worked].
R — Adopted department-wide; outlived my tenure. [Pass rates / feedback if you have them.]
This is an enablement story: you didn't just teach well — you built an asset other teachers scaled with. That's exactly how a partner org (and Bussa's team) thinks.
▶11 · The union stand — Dawson · fill the specifics
S — Executive council member, Teachers' Union + Science Program Committee — representing colleagues to the administration.
A — [One concrete position you argued against the administration or department: what was at stake, how you made the case, how it resolved — and if it went against you, how you carried the decision anyway.]
Second Backbone story (after #4) — five interviewers means the big LPs get asked more than once.
▶12 · The emergent-misalignment paper — research
S — Since April 2026: grant-funded (Coefficient Giving) independent research on emergent misalignment in LLMs — with two co-authors, no institution behind us.
T — Produce publishable, rigorous results on a moving research frontier.
A — Co-authored "Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs" — [your specific contribution: the experiments? the finetuning infra? the analysis? — say what YOUR hands did].
R — Accepted to two ICML 2026 workshops (AI4GOOD, Pluralistic Alignment); under review at NeurIPS 2026.
For Tapia/Bland: this is also a genAI-credibility story — you finetune and evaluate models yourself; "evals not vibes" is literally your day job.
▶13 · The methodology disagreement — research · fill the specifics
S — Research with co-authors means methodology fights with people you respect and no boss to settle them.
A — [A real disagreement with Rob/co-authors — experimental design, analysis choice, framing: what you argued, what the data eventually said, and how you committed if it went the other way.]
Ideal if this is the one that went against you — pairs with #4 where you may have prevailed. Committing gracefully to a peer's call is a stronger signal than winning.
▶14 · Critiquing a published metric — research · fill the specifics
S — A published AI paper ("The Hot Mess of AI") built its argument on an incoherence metric I believed was flawed.
A — Rather than just disagreeing, I dug into the metric's construction and wrote a public critique — [the core flaw in one sentence a non-specialist gets; what your analysis showed; what you'd measure instead].
R — [Status: published where? any response?]
For George: this is his "going beyond vibes" blog post in story form — evaluating AI claims with rigor instead of accepting them.
▶15 · AWS from zero → porting my own product — happening right now
S — Zero AWS experience when this process started in July. Passed the SA assessment in about a week by building my own study tooling.
A — Then I made it real: I run a production personal-assistant app (LLM agents, messaging channels, scheduling — one live user) on a bare VPS, and I'm migrating it to AWS right now, as infrastructure-as-code. Phase 0 shipped this week: CDK in Python — billing guardrails and cost alarms before the first dollar, CloudTrail, GitHub OIDC so CI gets roles instead of long-lived keys.
R — A real, phased strangler-fig migration plan — model supply to Bedrock, compute, storage, per-seam cutovers with rollbacks — written and being executed. Cost visibility landed before the first workload, which I've since learned is exactly the Well-Architected order.
Every interviewer knows you're new to AWS. This flips the weakness: "I don't just study AWS — I'm betting my own product on it, and I can tell you exactly where its services fit a real stateful workload and where they don't." Details if probed: chose Bedrock for model supply (the CLI supports it natively), kept SQLite workloads off network filesystems, Postgres as the path off-box.
Leadership Principles
All 16, mapped to the bank. ★ = heaviest-probed for SA loops. With five interviewers, expect the ★ ones twice — primary and backup listed.
How an LP answer gets scored
- STAR shape — real situation, your task, what YOU did, measured result.
- Specificity — names, numbers, dates. Vague = invented, in the interviewer's mind.
- "I" — your decisions, your hands. "We" answers score as observer.
- Self-awareness — what you'd do differently. Self-critical scores.
- The peel — follow-ups are the method, not suspicion. Welcome them.
- 60–90 seconds then stop: "happy to go deeper on any part."
The 16, mapped
- ★ Customer ObsessionStart with the customer, work backwards.→ #5 sellers · #8 estimators · backup: #10 teaching
- ★ OwnershipAct for the whole company, long-term. Never "not my job."→ #7 solo PoC · backup: #1, #11
- ★ Invent and SimplifyExpect invention; always simplify.→ #2 abstraction layer · backup: #7 vector insight, #10
- Are Right, A LotStrong judgment; seek to disconfirm your beliefs.→ #14 critique · #7 vector bet · mind-changed: #13
- ★ Learn and Be CuriousNever done learning.→ #15 AWS/porting · backup: #9 full-stack ramp
- Hire and Develop the BestRaise the bar; develop others.→ #10 teaching/mentoring
- Insist on the Highest StandardsRelentlessly high bars; fix problems so they stay fixed.→ #10 materials · #3 threshold controller only if the deck wasn't in this room
- Think BigBold direction that inspires results.→ #2 platform bet · career-scale bet on AI safety
- ★ Bias for ActionSpeed matters; most decisions are reversible.→ #7 PoC · backup: #15 assessment sprint
- FrugalityDo more with less.→ #2 cost routing · #7 vectors over GPUs · #15 budgets-first
- ★ Earn TrustListen, speak candidly, be self-critical.→ #11 union · #6 failure told well · #8
- ★ Dive DeepDetails, data, skepticism when metrics and anecdotes differ.→ #1 speedup · #12 paper · #14
- ★ Have Backbone; Disagree and CommitChallenge respectfully; commit fully.→ #4 architecture · #11 union · #13 co-authors
- ★ Deliver ResultsRight inputs, right quality, on time.→ #1 · #12 paper shipped · #3 only if the deck wasn't in this room
- Strive to be Earth's Best EmployerSafer, more productive, more just workplace.→ #11 union council
- Success and Scale Bring Broad ResponsibilityImpact beyond the company; be humble.→ #12 — AI safety is literally this
Follow-up survival kit
- "Why that approach?" → name the alternative you rejected and why.
- "How exactly?" → one level of real mechanism, then offer more.
- "What was the metric?" → the number if you tracked it; otherwise "I didn't instrument that — what I did measure was…"
- "What would you do differently?" → always have one.
- "Tell me about another time…" → that's why there are 15 stories, not 8.
- At your knowledge edge: "I don't want to guess — here's how I'd find out." That IS the SA answer.
The presentation
The loop includes a formal Technical Communications evaluation — retrieved from Amazon's official brief (PDF in the project folder). Your deck is built and deployed. One session hosts it; confirm which with Casey.
Their prompt, verbatim
"Tell us about a business problem that you solved with a technical solution. Describe the problem, the technical and business needs, and the technology used to solve the problem. Explain the tradeoffs you made and why the solution was the best one to meet the customer's needs."
- 3–5 slides with architecture or relevant diagrams
- 20 min presenting + 10 min Q&A interspersed (~30 total)
- Audience: mix of technical and non-technical leaders — adapt messaging both ways
- Doesn't need AWS · nothing confidential · they expect only 1–2 hours of prep
Your deck: "One Interface, Many Models"
Open the deck →Two Weeks to Kickoff — the regional threshold controller for Euro 2024, structured exactly to their rubric: problem → business + technical needs → options and tradeoffs → architecture diagram → impact with honest limitations. (The earlier LLM-layer deck is kept at /deck-llm-layer.html.) Keys: ← → navigate · N presenter notes · F fullscreen. Notes carry your timing plan and the [red gaps] to fill.
Timing: 1' title → 4' problem/needs → 6' architecture → 5' tradeoffs → 3' impact → Q&A buffer. Q&A is "interspersed" — welcome interruptions, they're the point.
What they're scoring (from the brief)
- Understanding of the business requirements + impact of your solution
- Conveying technical concepts to both technical and non-technical roles
- Accuracy of the technical content
- Clear, logically-flowing answers to questions
- Adapting messaging to the audience — watch faces, offer the plain version
- Articulating tradeoffs, benefits, risks, limitations — slide 4 is dedicated to this
Deck prep checklist
The Bedrock parallel is your closing line, not a slide — it lands harder spoken: "Bedrock is this idea as a managed service; I built the small version, so I know why it matters."
Question drill
Real-style questions for the loop. Answer aloud in 90 seconds, reveal the intended story + what a strong answer includes. "Got it" retires the question on this device.
Loading…
Suggested schedule (Mon–Wed)
- Mon: fill the seven [red-gap] stories (Voice #2), then 6–8 drill questions.
- Tue: full mock loop (Voice #1) + presentation rehearsal (Voice #3). Sabri session if scheduled.
- Wed: 8–10 drill questions on ★ LPs · dive-deep grill (Voice #4) · deck run-through · logistics checklist. Stop by dinner.
- Thu 9:30: morning warm-up (Voice #6) only. No cramming.
Voice practice
Copy a prompt → Claude app → new chat → paste & send → voice mode. Each is self-contained.
Which one, when
- Mon: #2 Gap-story workshop (builds the seven [red-gap] stories).
- Tue: #1 Full mock loop · #3 Presentation rehearsal.
- Wed: #4 Dive-deep grill · #5 Rapid-fire recall.
- Thu morning: #6 warm-up. Nothing else.
1 · Full mock loop interview (~40 min)
Rotates interviewer personas from your real panel — AE plain-language round, partner-manager round, technical round, hiring-manager round — with scored debrief.
2 · Gap-story workshop (~30 min)
Builds the seven unfinished stories by interviewing you, then stress-testing each.
3 · Presentation rehearsal (~35 min)
You deliver the 20-minute deck aloud; Claude plays a mixed technical/non-technical audience, interrupts with real questions, then scores you against Amazon's actual rubric.
4 · Dive-deep grill (~15 min)
A skeptical Principal drills one technical story five levels down; teaches the honest knowledge-edge move.
5 · Rapid-fire recall (~15 min)
All 15 stories, hook + number in one breath each, then quick LP call-outs. Cardio for retrieval.
6 · Morning warm-up (~15 min · Thursday only)
Light reps, high energy, ends with logistics reminders. Leaves you warmer, not drained.