7th Ann. MLOps World | GenAI Conference & Expo
About this Event
👋 Thank you for joining us here
The goal of the MLOps World | GenAI Summit 2026 is to help companies move from experimentation to production, where AI Agents and Agentic Workflows are transforming how teams build, scale, and operate AI.
This year’s flagship theme is AI Agents & Agentic Workforces, with a focus on:
- AI Agents for Developer Productivity: accelerating workflows and innovation
- AI Agents for Model Validation & Deployments: ensuring reliable systems in production
- Augmenting Agentic Workforces: scaling human + AI collaboration
- Agents in Production: lessons from real-world deployments
- Latest Trends in MLOps: staying ahead with evolving best practices
Tracks:
- AI Sovereignty
- AI-Assisted Software Engineering
- Agent Harness Engineering
- Evals and Benchmarks
- Cost Management and ROI
- Agent Memory Architectures
- Recursive Self-Improvement
- Hardware and Chips
We’re a group of practitioners committed to sharing case studies, architectures, and proven strategies, no hype, no product pitches.
Just hard-earned lessons from teams deploying agents and agentic systems at scale.
Whether you’re exploring copilots, scaling agent stacks, or running multi-agent orchestration, this summit is designed to give you the insights and tools to succeed in production.
👉 Join us in Austin this October to connect with over 700+ AI engineers, platform teams, and leaders building the future of AI Agents in production.
- MLOps World Team
General Admission ticket include;
✅ Complete access to the Summit Day 1 (Nov 17th) and Day 2 (Nov 18th)
✅ Complete access to the bonus virtual talks and workshops on Nov 16th
✅ Access in-person talks, networking with food and drinks
✅ AI-Powered Desktop & mobile event app for online networking
✅ Conference parties
✅ Access to post-summit videos
MLOps World is an international community group of practitioners working to advance the science of deploying not only ML models but also agents and agentic workflows into live production environments. We explore everything, technical and non-technical, that goes into making these systems reliable and effective.
With an explorative approach, our initiatives address the needs of a community of over 18,000+ ML and AI practitioners, researchers, professionals, entrepreneurs, and engineers. We’re here to empower members, propel productionized AI, and shape the next generation of agent-driven systems.
Our gatherings and events aim to reimagine what it means to have a connected community, offering support, growth, and inclusion for all participants.
Nov 16 - Virtual Session
🕑: 10:20 AM - 10:50 AM
Scaling AgentOps: Observability, Safety and Control in Production AI
Host: Andy McMahon, Principal AI and MLOps Engineer, Barclays
Info: As organisations move from single models to connected agent systems, the challenge shifts to how agents interact, how they are evaluated, and how their behaviour can be monitored in real production environments.
This presentation explores what it takes to operationalise agentic AI in practice – from observability and evaluation to deploying agents safely and reliably across enterprise workflows.
Through real-world examples, we look at how teams are building responsible AgentOps frameworks to scale multi-agent systems in production while maintaining control, trust and measurable business impact.
🕑: 10:55 AM - 11:25 AM
Placeholder
Host: Yegor Denisov-Blanch, Research Scientist, Stanford Unive
Info: Placeholder
🕑: 11:30 AM - 12:00 PM
Beyond Unit Tests: A digital-Twin Approach to AI Agent Evaluation
Host: Vicente Ruben Del Pino Ruiz, Sr Director AI & Data Engin
Info: Single-shot evaluation, the shape every current AI eval framework ships, is structurally blind to the failures that take agents down in production: patience that runs out at turn six, users who abandon silently, tool calls that lose conversation state, partial successes that masquerade as wins. This talk argues for a different shape, drawn from the digital twin literature and how the approach is already applied in aerospace, autonomous vehicles, and civil engineering: simulate the population of users your agent will meet, run them against the agent, watch what breaks before any human sees it.
Three take aways:
1. Why current AI agent evaluation cannot detect the failures that hit production. The structural reason prompt-and-grade testing (the shape every current framework ships) is blind to patience, abandonment, conversation state, and partial success.
2. What a realistic synthetic user population looks like, drawn from the digital twin literature in other engineering industries.
🕑: 11:30 AM - 12:00 PM
There's No One There to Press Retry: What We Learned Running AI on a Million D
Host: Antonio Bustamante, Cofounder & CEO, BEM AI
Info: Everyone building on frontier models hits the same wall: 80% of the way there in a weekend, then an exponentially expensive climb toward the 99%+ that operational systems need. This talk is the honest map of that climb, from a team that now runs AI over millions of documents, images and videos a month for customers in logistics, fleet management, automotive and financial services who need the answer to be right every time.
We will walk through the ladder nobody budgets for: retries, then queues when the model is down for three hours, then rate limits, then discovering that 95% is not enough for transactional data, then discovering that the model cannot tell you how confident it is. We will show two failures from our own production history, a customer whose single engineer racked up $30K of usage in a month because nothing was watching, and our first churn, on 500-page reports with a hundred rows per page, a problem we still consider unsolved. And we will show what we changed: a harne
🕑: 12:05 PM - 12:35 PM
Silent Drift: Why Your LLM's Quality Is Degrading and Your Metrics Can't See I
Host: Sandeep Bharadwaj Mannapur, Lead AIML Engineer
Info: We caught our LLM quality problem the wrong way: through customer complaints. Three weeks of degraded responses had already shipped. Every monitoring signal we had was green the entire time because we were measuring the wrong things. Latency, error rate, embedding similarity: none of them captured what was actually happening, which was that response quality had quietly gotten worse in ways users noticed but our dashboards could not.
After that incident we rebuilt how we think about LLM observability. The core insight was that output quality is not the same as system health, and you cannot infer one from the other. We needed a separate signal layer built from production behavior: how users responded to answers, where downstream tasks broke down, where retrieval and response stopped agreeing with each other.
Getting this right took a few tries. A naive implementation fires constantly on normal LLM output variance. The real work was designing alerts that distinguish actual degrad
🕑: 12:40 PM - 01:10 PM
Fan-Out, Tail Latency, and Cluster Mutations: Three Problems in an Ad-Serving
Host: Jim Allen Wallace, Product Marketing, Dragonfly
Info: "Instacart's ad-serving feature store served features to real-time inference from a multi-hundred-node managed Valkey deployment split across multiple clusters. The team had built a proxy layer, a querying SDK, and a compact storage format on top of it, and still hit three limits: instability during any cluster mutation, tail latency from read fan-out, and cost that got worse when they split clusters to manage the first two.
The structural change was to stop scaling out and scale up. A multi-threaded engine let the team run much larger instances, so hundreds of nodes became roughly 100, each inference request touched far fewer shards, and average and P99 latency fell 50%. Because the proxy and SDK already hid the datastore from ML engineers, the swap required no application changes and tens of terabytes moved without anyone above the storage layer noticing.
We were the engine vendor, and the first pass did not reach the latency the team expected. Getting there took client-side tuning
🕑: 12:40 PM - 01:10 PM
Closing the Gap Between What AI Is Trained On and What Users Actually Need
Host: Jazmia Henry, Research Scientist University of Oxford
Info: There's a gap between researcher-crafted evaluation frameworks that capture model performance and benchmarks versus how end users actually use AI products. This gap is being exploited in ways that render traditional reward models useless.
🕑: 01:15 PM - 01:45 PM
Building Agentic Apps with Customer Data Products
Host: Matt Mazzarell, AI Lead, Financial Services, Americas Te
Info: "One of the most difficult problems every company faces is understanding its customers completely. Customer lifetime value, attrition risk, and purchase propensity are all solvable with AI/ML — but how do we combine these modeling scores to initiate the right action with the right customer at any point in time?
Agentic applications help us make the best possible decisions when interpreting complex, high-volume signals from our customers. An agentic application gives end users visuals that explain key insights, with an agent in the loop to ensure nothing is missed. Context is everything: when done correctly, the agent always has the appropriate understanding to build an action plan that improves customer health and profitability.
In this session, we'll show you how to build agentic apps from ideation to a finished product that interacts with customers. You'll take away practical tips for using agentic coding frameworks, curating complete customer data products, and building customer-f
🕑: 03:00 PM - 03:30 PM
From Python Function to Serverless GPU with Runpod Flash
Host: Jessica Garson Beauchemin, Developer Relations Lead Runp
Info: Deploying AI applications often involves containers, GPU configuration, dependency management, and infrastructure scaling. Runpod Flash offers a simpler, code-first approach. It lets you define remote functions and hardware requirements in Python and run them on serverless GPUs.
In this talk, we’ll explore how Flash moves Python workloads from local development to cloud deployment, examine its underlying programming model, and build a GPU-backed endpoint. Along the way, we’ll discuss where Flash fits in the AI development stack, the problems it solves, and the trade-offs developers should consider. Attendees will leave with a practical understanding of how to turn local AI code into a scalable service without needing to become infrastructure experts.
Nov 17
🕑: 09:45 AM - 10:25 AM
Keynote: Maybe the Puppets Were Right All Along
Host: D. Sculley, Former CEO, Kaggle
Info: We spend a lot of time thinking about operational issues in AI related to deployment, somewhat less time thinking about adoption, or (dare we say it) acceptance. This talk will touch on some technical pieces in the current AI ops landscape including streaming systems, planning, latency, and on-device models, but most of the time will be spent looking at ways we can move beyond the stale framing of a chatbot, assistant, or customer service agent.
🕑: 10:55 AM - 11:25 AM
Driving Better Outcomes with Open Weight Models
Host: Anne Griffin, Founder & AI Product Consultant Griffin Pr
Info: When Hugging Face went to investigate this summer's OpenAI containment breach, the commercial models they reached for refused to help, because their safety guardrails couldn't tell defending from attacking apart. Instead, Hugging Face used an open weight model on their own infrastructure and finished the forensics in hours with none of the attacker data leaving their environment. It's the perfect example of how open weight models let you control your guardrails, privacy, governance, latency, and reliability.
More companies are asking if they should use and self host open weight models, and this talk will walk through a framework to determine when it does and doesn't make sense. By the end of this talk, attendees will understand what the benefits and drawbacks are of open weight models, and how they can bring both technical and business value.
🕑: 10:55 AM - 11:25 AM
AI-Assisted, Audit-Ready: Scaling ML Platform and Developer Productivity at a
Host: Manikandan Paramasivan, Principal Architect - Data, ML a
Info: Building ML models is not the hard part for most teams anymore. Getting them from a data scientist's notebook to a governed, monitored, production system and keeping them there without a platform team standing between every developer and every deploy is. It gets harder inside a regulated fintech, where every model decision needs an audit trail, every alert needs a documented escalation path, and every """"quick fix"""" has to survive a compliance review.
This talk walks through how our ML platform team at KOHO, a Canadian fintech, built a developer-first MLOps stack on AWS that treats governance as a byproduct of good engineering rather than a tax on top of it and how we're now layering AI-assisted workflows (Claude Code skills, agentic guardrails) on top to remove the remaining manual toil from model onboarding, monitoring triage, and incident response.
We'll cover the architecture, the AI-assisted developer tooling built on top of it, and what changed operationally as the compa
🕑: 10:55 AM - 11:25 AM
Letting AI Agents Run Incident Response on a 12M-Device Network: Guardrails, G
Host: Arun Malik, Principal Software Engineer, Azure Networkin
Info: We redesigned frontline incident response around AI agents that do not just suggest fixes but execute them, across more than 12 million devices and tens of thousands of incidents a month. Structurally, three things changed. First, agents stopped getting broad human-equivalent access and instead call tools through a governed interface, with permissions scoped down to individual functions and parameter values, so a compromised or confused agent has a bounded blast radius. Second, human approval moved from a blanket gate to a targeted one, applied only to sensitive or irreversible actions and routed by how familiar the problem is and how reversible the action is. Third, we stopped paying full model-inference cost for repeated work by promoting an agent's proven, validated behavior into deterministic playbooks that run at near-zero token cost, which cut agent running cost by more than 70 percent over eight months while incident volume doubled.
Attendees will leave able to decide what
🕑: 11:30 AM - 12:00 PM
Building Reliable Browser Agents in Healthcare
Host: Jake Kang, Co-Founder Whitney AI
Info: "This talk shares lessons from building a self-recovering browser agent for regulated healthcare workflows, where we moved from a single-agent prototype to a multi-step system with optimized inputs, memory, evals, sandboxes, and deterministic code where possible. I’ll walk through how we used production failures to build a faster evaluation and prompt optimization loop, and when to move beyond harness engineering to increasing model capability via post-training.
What You’ll Learn
Attendees will learn components of the agent development lifecycle and how to build an effective harness, and when to fine-tune when you hit a plateau with foundation models:
- how to turn production traces into eval datasets
- how to manage context and memory
- how to build sandboxes that speed up testing and iteration.
🕑: 12:05 PM - 12:35 PM
From Model Metrics to System Behaviour: Measuring Reliability in Non-Determini
Host: Aryan Dhar, Senior MLE, Wisedocs
Info: This talk presents a practical framework for evaluating increasingly complex, non-deterministic AI systems, moving from individual models to complete production pipelines. I begin with a page-stream segmentation model, where standard metrics concealed important failure modes at Wisedocs. For example, we discovered that the accuracy of identifying a document boundary is not the same as the ability to reconstruct the complete document span. Furthemore, errors on some document types carried greater downstream risk than others, and small aggregate improvements did not necessarily translate into meaningful business outcomes. This case study shows how evaluation for us evolved towards task-specific metrics, sensitive failure slices, strict regression testing, and more business-facing metrics of reliability.
I then move to LLM-based systems using Medical Long-Context Reasoning, or MLCR, as a case study. This was a benchmark we developed at Wisedocs earlier this summer. I will examine how
🕑: 12:05 PM - 12:35 PM
Review without Skew
Host: Robert Lewis, Senior AI engineer Precocity, LLC
Info: This talk presents a practical framework for evaluating increasingly complex, non-deterministic AI systems, moving from individual models to complete production pipelines. I begin with a page-stream segmentation model, where standard metrics concealed important failure modes at Wisedocs. For example, we discovered that the accuracy of identifying a document boundary is not the same as the ability to reconstruct the complete document span. Furthemore, errors on some document types carried greater downstream risk than others, and small aggregate improvements did not necessarily translate into meaningful business outcomes. This case study shows how evaluation for us evolved towards task-specific metrics, sensitive failure slices, strict regression testing, and more business-facing metrics of reliability.
I then move to LLM-based systems using Medical Long-Context Reasoning, or MLCR, as a case study. This was a benchmark we developed at Wisedocs earlier this summer. I will examine how
🕑: 01:35 AM - 02:05 AM
I Made 104 Model Configurations Play Mafia Against Each Other
Host: Drew Crawford, Owner
Info: Somewhere in a datacenter right now, one frontier model is telling another that it's the town doctor. It is not the town doctor. It killed someone four turns ago, and it's about to get a third model lynched for it.
Most LLM benchmarks are exams: a model alone in a room with a test paper. Mafia is a room where some of the agents are lying, everyone knows some of them are lying, and the game is figuring out which. It demands recursive theory of mind, deception that stays consistent under adversarial re-reading, and long-context discipline — and it punishes output indiscipline like no static eval can: a model that rambles or breaks format doesn't get a bad grade, it gets voted out. The other players are the evaluation harness.
This talk covers what it took to run AI Bot Mafia in production: 40,000 lines of Rust; the streaming pathologies of 104 model configurations (chunks that split UTF-8 characters mid-byte, keepalive-only hangs no single timeout catches, malformed tool calls that
🕑: 02:10 PM - 02:40 PM
When Pricing, Catalog, and Compliance Collide: Building Multi-Agent AI Swarms
Host: Amit Kumar Padhy, Senior Computer Scientist II & Lead Ar
Info: Modern commerce platforms don't fail because of missing features, they fail at the seams.
A product is created in Catalog, but pricing is incomplete. Promotions don't qualify. Tax blocks specific regions. Localization lags. The system says """"launched,"""" but the business knows it isn't. These aren't edge cases, they are the steady state of distributed commerce.
This session replaces traditional workflow orchestration with a multi-agent, swarm-based execution model powered by LLMs and agentic AI, coordinating Pricing, Catalog, Promotions, Tax, and Compliance in real time.
Agent Roles. Planner Agents decompose onboarding goals into executable plans using ReAct-style tool-aware reasoning and function calling. Domain Agents for Pricing, Catalog, and Compliance execute directly against APIs, Pricing Runtime, Offer Systems, Billing Preview, Tax engines. Validator Agents enforce policy rules, regional compliance, and pricing integrity at every step. A Coordinator Agent maintains shared
🕑: 02:45 PM - 03:15 PM
When the Bill Arrives: Cost Engineering for Production AI Agents in Enterprise
Host: Vishvesh Pandey, Quant Analytics Senior Associate JPMORG
Info: Most practitioners instrumenting AI agents in production analytics workflows start by tracking token consumption — and discover too late that token spend tells you what you spent, not whether it was worth it. This talk walks through the operational reframe underway in financial services and adjacent regulated industries: from cost-per-token to cost-per-resolved-task as the primary economic unit, with call-level attribution as the foundation that makes everything else possible.
I'll cover four cost engineering patterns that hold up in production analytics environments: (1) instrumentation at the call level — what to log, where it lives, who owns it; (2) model tiering and routing — when a frontier model earns its premium and when a smaller fine-tuned model wins; (3) the output-token asymmetry that makes structured outputs and chain-of-thought hygiene more economically consequential than they appear; and (4) cost containment as an engineering practice with runtime guardrails, not a fi
🕑: 02:45 PM - 03:15 PM
The Log Is the Agent
Host: Ishaan Sehgal, CEO Omnara (YC S25)
Info: At Omnara, long-running agent sessions exposed a failure mode hidden by short-lived agents: the agent’s state was coupled to the process, model provider, or runtime executing it. When a worker crashed, a client disconnected, or execution moved to another machine, recovering the agent’s work became brittle.
We redesigned agent execution around an append-only session log as the source of truth. Model outputs, tool calls and results, user interventions, and other state transitions are persisted as ordered events; the live agent state becomes a replayable projection of that history.
This structural change decouples agent identity from any one model, worker, or harness. It enables crash recovery, resumability, branching, real-time observability, and migration across models and machines, but introduces difficult systems questions around ordering, idempotency, partial tool execution, checkpointing, compaction, concurrent writers, and log ownership.
This session walks through the product
🕑: 03:45 PM - 04:15 PM
Why Your Multi-Agent Pipeline Is Slow and Expensive (And How to Fix It Systema
Host: Nadia Rauch, AVP, AI Engineering | Principal AI Engineer
Info: We built a multi-agent pipeline to automate a complex, document-intensive enterprise workflow. The initial system ran sequentially through a chain of specialized agent roles, each handling a distinct task in the process, using a single frontier model throughout. It worked. It was also slow and expensive, and we didn't know why.
Rather than optimize blindly, we profiled first. The results were not where we expected: 67% of total latency came from a small minority of the agent roles, and the primary bottleneck was not the model — it was sequential chaining: independent work was being processed one step at a time instead of concurrently. Identifying and parallelizing the roles with no inter-dependency reduced end-to-end runtime by roughly 40%.*
The second experiment compared model tiers (frontier, mid-tier, lightweight) across each agent role independently, measuring accuracy, latency, and cost per role. The finding cuts against the default assumption: the roles that appea
🕑: 03:45 PM - 04:15 PM
Personalized Experiential Learning: How Do Enterprises Build AI Agents That Le
Host: Kumaran Ponnambalam, Principal ML Engineer Cisco Systems
Info: This session examines how a traditionally static enterprise agent workflow was redesigned around AI-driven adaptation. In most enterprise settings, agent behavior is defined through fixed prompts, rules, and workflows that apply broadly across users and tenants. We redesigned that model into an adaptive agent architecture built around explicit policy layers, feedback loops, and governed personalization. Structurally, this changed the system from one-time configuration to continuous learning: policies were separated from model reasoning, policy scope was organized across global, tenant, and user levels, and agent behavior was connected to both explicit and implicit feedback signals that could influence future decisions.
The session will focus on what changed in the system architecture and operating model to support this shift. It will cover how policy-driven personalization was introduced, how feedback became part of the runtime and improvement loop, and how adaptation was managed
Build the Pond Before You Teach Fishing: Killing Per-Person AI Tooling and Mea
Host: Tony Blank, Staff AI Engineer
Info: We stopped equipping humans and started automating their worst hour. Instead of teaching people to fish, we built the pond: precise background agents pointed at repeated, skill-independent time-sinks where ROI was measurable on day one. The first build was a bug-triage agent wired to GitHub, Jira, and Notion; every agent that followed was measured the same way — Time Saved Per Task (TSPT) per week — against an amortized cost model (build + inference + maintenance vs. hours saved × loaded hourly cost).
Structurally, the change was moving from a portfolio of per-person tool subscriptions to a portfolio of scoped, measurable background agents with a pre-build attribution conversation baked in. The decision rubric — Sharpness × Cadence × Skill Tax → Pond / Boat / Net / Wade — tells you when to automate fully, when to augment, when to keep a human in the loop, and when the math doesn't work yet.
Attendees will leave able to: score their own candidate workflows on that rubric, instrume
Nov 18
🕑: 10:55 AM - 11:25 AM
The Model Was Right. The Corrections Still Failed.
Host: Deji Andrew, Manager, Systems & Data Platforms Niagara B
Info: A LightGBM model was deployed across production lines to predict and correct overruns mid-job, reducing average overproduction from +24 cases to +2 on historical data. In pilot, the corrections worked: overruns vanished. But a closer look at the production data showed something different. When the order target was reduced by 21 cases, the line overproduced the new target by 18. The system was not being corrected - it was negotiating with the correction, regenerating overrun relative to whatever target the machine was now working toward. Compounding this: corrected jobs left no trace in the training data that a human had intervened, meaning the model, on retraining, would learn to stop flagging the problem it had been built to solve. This talk walks through the pilot data, explains why conservatism in the loss function wasn't enough, and argues that the real gap isn't in the model - it's in the experiment infrastructure that doesn't exist between a prediction and a PLC.
🕑: 10:55 AM - 11:25 AM
One Orchestration Layer for All: Unifying Batch ML, Real-Time Inference
Host: Maitrik Patel, Sr Engineering Manager, Apple
Info: Running ML training, AI inference, and agent workflows on three separate orchestration systems is an operational debt that compounds with every new workload type - one system per abstraction means three on-call rotations, three observability stacks, and three onboarding paths. We collapsed this to a single execution layer by designing a type-safe task interface that accommodates non-deterministic agent execution alongside deterministic ML pipelines, backed by container-native isolation and a workload-aware shared scheduler. The unification exposed a hard failure class that static ML infrastructure never encounters: agent tasks are not idempotent, and ML-derived retry semantics caused downstream state corruption in early production until we rebuilt retry logic around explicit checkpointing and action journals. Attendees leave with a four-layer architectural blueprint for unifying these workload types, an honest account of the scheduling contention and retry failures encountered in prod
🕑: 10:55 AM - 11:25 AM
Performance Is the Product: Our Journey to SoTA for Nemotron Retriever
Host: Kalpesh Sutaria, Director NVIDIA
Info: A leaderboard win is not a product. Nemotron Retriever's models rank #1 on RTEB — but topping a benchmark and running efficiently inside a customer's production are two very different problems. This talk is the engineering story of closing that gap.
I'll walk through our journey rebuilding the Retriever inference stack in Rust with a single obsession: treating state-of-the-art performance as a first-class product requirement, not a post-hoc optimization. We'll get concrete about the decisions that mattered — what we measured, where we spent effort, the throughput, memory, and footprint wins we chased, and how dramatic efficiency gains unlocked deployment scenarios (including edge and on-device) that simply weren't possible before. I'll also share what didn't work, the tension between shipping fast and building durable, and how we leaned on AI-assisted development to move faster than the roadmap assumed.
You'll leave with a mental model for treating inference performance as a prod
🕑: 11:30 AM - 12:00 PM
The Economics of Autonomous Research
Host: Vashishtha Patil, Senior Applied Scientist
Info: Autonomous research agents that propose, implement, and refine ML solutions are priced out by the loop, not the model. Published agents assume a frontier model drives every step, so cost scales with trajectory length on runs that are long by design. Teams either cap the horizon — removing the thing that made it work — or don't run it at all.
Hypothesis: a small open-weight model runs the loop end to end, with a frontier advisor called in at those points on a metered budget, holding most of the baseline's result quality at a fraction of its cost. The open question is whether the escalation trigger is reliable enough to justify the calls it buys — our first version often bought advice the loop would have reached on its own.
🕑: 11:30 AM - 12:00 PM
Maximizing GPU Utilization in LLM Post-Training with Co-Operative Time-Slicing
Host: Poonam Lamba, Senior Product Lead, Google
Info: We redesigned distributed GPU orchestration for RL post-training and batch inference in the open-source llm-d platform. Structurally, we replaced static GPU/TPU locking with a three-tier co-operative time-slicing system:
Application Layer: Workloads signal phase boundaries (rollouts, training, batch inference) via explicit acquire() and yield() APIs.
Cluster Orchestrator: Manages lock queues to dynamically interleave complementary jobs onto shared hardware during idle phases.
Node Snapshot Agent: Executes fast sub-second state swaps between GPU/TPU VRAM and host DRAM, enabling instant context switching without container restarts.
Attendees will walk away with:
Drive 70%+ GPU/TPU Utilization: Understand how time-slicing reclaims idle hardware during RL loops and batch inference without impacting convergence.
Architect Rapid Memory Swapping: Apply VRAM-to-DRAM snapshotting strategies for ultra-fast GPU/TPU context switching.
Deploy on Kubernetes: Configure llm-d and K8s orchestrator
🕑: 12:05 PM - 12:35 PM
Closing the Loop in the Agentic Software Development Lifecycle with Evals
Host: Zachary Hamilton, Solutions Engineer, Braintrust Data
Info: Building agentic software requires more than adding an LLM to the traditional software development lifecycle. When behavior becomes non-deterministic, teams need a new feedback loop connecting what happens in production to how they evaluate, debug, and improve their systems. This talk will show how to build that loop: identifying real failure modes from production traces and human feedback, turning them into reproducible datasets and evaluations, defining success using both technical and business outcomes, and using automated and human judges to measure whether changes actually improve the system. We'll also explore where automated evaluation breaks down, how to calibrate LLM-as-a-judge, and how continuous evaluation helps teams detect regressions and emerging behavior. Attendees will leave with a practical framework for moving from production > failure discovery > datasets > evals > iteration > production, turning the development of agentic systems into a measurable engineering discip
🕑: 12:05 PM - 12:35 PM
From Prompt to Privileged Action: Identity and Audit Controls for Enterprise A
Host: Siddharth Jain, AI Engineering Manager, OpenAI
Info: Teams often give an agent a service credential, add a human approval step, and call the workflow governed. That design breaks down when the agent can revise payloads, retry writes, chain tools, or act across systems with different permission models. The result is an accountability gap: the organization can see that a service account acted, but not necessarily who authorized the business intent, which payload was approved, whether a retry duplicated work, or what changed between proposal and execution.
This session presents a production control model for agent workflows that make consequential writes. It shows how to separate the business-intent identity from the concrete operation; classify tools by impact; issue short-lived, least-privilege credentials only after validation; bind human approval to a canonical payload hash and policy version; execute through controlled services with idempotency keys; and reconcile external state before declaring success. It also covers the evidence
🕑: 12:05 PM - 12:35 PM
Training Generative Recommenders in Production
Host: Fuzail Khan, Senior Machine Learning Engineer, Meta
Info: The ranking and recommendation systems landscape is being transformed in the generative era. This talk reports on the experience of building and shipping the training infrastructure behind the first generative recommender in production at Meta, covering both stages of the recipe: pre-training to acquire the generative capability and post-training RL to align generation with the ranking objective.
Generative recommenders bring distinct challenges to end-to-end performance and scalability. This primarily arises from a mixed architecture that consists of both recommender-native large sparse embedding tables and LLM-based decoders. This then generates semantic IDs translating to real-world use cases in production such as finding the right advertisement for a given user. We inherit the communication profile of a sparse recommender as well as the autoregressive nature of a large language model for which an end-to-end systems blueprint simply does not exist.
For pre-training, we make
🕑: 01:35 PM - 02:05 PM
Little's Law Meets the Roofline Model: A Decision Framework for Sizing Hosted
Host: Balaji Varadarajan, Staff Engineer, LLM Inference Digita
Info: We will start with two diagnostic tools: Little's Law which translates your QPS and target latency into the concurrency and batch size your system actually needs and the roofline model which tells you whether prefill (compute-bound) or decode (memory-bound) is your real constraint before you reach for a fix..
Will walk through the levers that move the needle most - MoE vs. dense architecture and what that does to your communication pattern FP8/INT4 quantization and where it costs you accuracy, TP/EP/DP parallelism and when each is the right tool and batching to the B_sat knee before compute-bound latency hockey-sticks..
Will cover concrete tuning playbooks for two workload types -
1. latency-sensitive traffic like chat and agents where you're optimizing TTFT and P99 ITL with chunked prefill and modest batch sizes
2. Throughput-sensitive traffic like batch summarization and evals where you push past B_sat and let queuing work in your favor.
Most production systems live in both
Nov 19 - Workshops
🕑: 11:15 AM - 12:45 PM
Make It Never Fail: A Hands-On Lab in Taking AI Extraction from 80% to Product
Host: Upal Saha, Co-founder & CTO
Info: This is a build-and-break lab for engineers who already know that a demo is not a system. In the first fifteen minutes every attendee stands up a working extraction pipeline against a real invoice, with no schema authoring, and gets structured JSON back. Then we spend an hour breaking it the way production does, and fixing each break with a pattern that transfers to any stack.
Break one: the wrong document. We feed a bill of lading glued to an invoice into the invoice pipeline and watch it confidently produce garbage. Fix: classify before you extract, and make the graph deterministic even though every step inside it is a model. Break two: the answer that is probably right. Models do not tell you how confident they are, so we compute per-field confidence, find the fields that fall below 95%, and route those, and only those, to a human whose correction feeds back into the system. Break three: the messy string. ""10 cases organic gala apples, 88 ct"" has to become one SKU; we show why c
🕑: 11:15 AM - 12:45 PM
Building Better Coding Agents: A Hands-On Workshop in Harness Engineering
Host: Rajiv Shah, AI Engineer OpenHands
Info: You can start a simple agent with a model and a prompt. But you soon realize that improving it requires adjusting what the model can see, what it can do, how it remembers progress, which model handles each step, and what evidence allows the work to stop. Harness engineering is how those pieces become one working system.
This workshop uses coding agents to make that system concrete. Participants will work through six decisions: choosing a harness, designing tools and retrieval, placing context and memory, routing models and reasoning, controlling goals and validation, and deciding when a task benefits from multiple agents. Through hands-on experiments, we will change these settings and inspect what happens in the trace and final result.
Each exercise is paired with an overview of current research on agent harnesses. Together, the research and experiments show how the major components interact and where each approach reaches its limits. Participants will leave with a practical way to
🕑: 01:00 PM - 02:30 PM
NemoClaw: Building a More Secure Runtime for Long-Running Autonomous Agents
Host: Chris Alexiuk, Sr. Product Research Engineer, NVIDIA
Info: In this workshop we'll stand up NemoClaw end to end: install the reference stack, get OpenClaw running inside the OpenShell sandbox, configure inference routing, and lock down a network policy that survives a multi-hour agent session. We'll walk through the blueprint, the CLI, and the approval flow, then run a real long-lived agent against it and break things on purpose so you know what the layers actually catch.
Where is it happening?
Event Location & Nearby Stays:
USD 0.00 to USD 675.46


















