#StayHappening
Thu, 05 Nov, 2026 at 08:30 am to Fri, 06 Nov, 2026 at 03:00 pm UTC-05:00

MLOps North: Building (with) Agents

  • RBC WaterPark Place · Toronto
  • From CAD 0.00
Toronto Machine Learning Society (TMLS) Publisher / Host Toronto Machine Learning Society (TMLS)
Share
MLOps North: Building  (with) Agents
Advertisement
A gathering for people building agents and the practitioners building with them.
About this Event

MLOps North Toronto Building with Agents

Toronto Summit · Presented by TMLS · Nov 5–6, 2026 · RBC WaterPark Place, TorontoThe people building agents, and the people building with them - one room, on the Toronto waterfront. Two days of practitioner-curated talks on what actually ships: the stack underneath agentic systems, and the products, workflows, and agents teams are running in production right now.

Why come

This isn't another "Intro to RAG" track. It's the things people are debugging, scaling, and fixing this quarter - from teams who've done it. Every session is peer-reviewed by working engineers, not a vendor keynote in disguise.

The theme is split into two tracks

Track 1 — Agents in Production

  • Agents in Production
  • Agent Infrastructure
  • Systems & Reliability
  • Deploying Agents at Scale
  • Observability & Evals
  • Scaling Agentic Systems

Track 2 — Agentic Development

  • Agentic Development
  • The Agentic Engineering Workflow
  • Developer Productivity with Agents
  • Engineering with AI Agents
  • The AI-Native Dev Workflow
  • Tooling & IDE Integration
  • Team Adoption & Governance

What's inside

  • 2 days, in person
  • 10+ tracks
  • 35+ speakers
  • Talks, hands-on sessions, and hallway conversations with the people running this in production

Who it's for

Software engineers, ML and data engineers, solution architects, infra leads, and the technical leaders (Director, VP, C-suite, founder) making the calls on how AI gets built and shipped.

Details

Nov 5–6, 2026 · RBC WaterPark Place, Toronto waterfront.


NOV 5

🕑: 10:55 AM - 11:40 AM
Agentic ML - Leveraging Harnesses to Automate Extraction Models
Host: Aryan Dear, Senior ML Engineer

Info: Model, Data and Objective Drift have traditionally ML Engineering problems that have had a ceiling where we could automate. This talk is a case study of how Wisedocs automated large portions of model development across a 1500+ document types that face standard MLOps challenges. The talk will have three parts, a focus on domain knowledge, the ROI of agentic ML development and sample harnesses and skills we use to enable these capabilities.


🕑: 10:55 AM - 11:40 AM
From Answers to Evidence: Building Auditable Agents
Host: Akash Shetty, CTO, Publicus

Info: Publicus processes procurement data spread across RFPs, contracts, amendments, award notices, invoices, task authorizations, webpages, and other records. Our early AI workflows optimized for producing a useful answer: retrieve relevant context, reason over it, generate structured output, and evaluate the result. That worked until we tried to increase automation in workflows where an answer can be semantically correct while its evidence is incomplete, stale, contradictory, or simply the wrong source.

We redesigned the system around an evidence ledger rather than the generated answer. Extracted facts and agent claims became typed artifacts carrying source document, location, provenance, transformation history, confidence, and conflicts. Retrieval became evidence acquisition rather than context stuffing. Verification became a separate stage that reconciles claims across records. Unsupported or conflicting outputs cross an explicit human-review boundary instead of being silently conver


🕑: 11:45 AM - 12:30 PM
Loop Engineering is just K8s-style Cybernetics
Host: Sayantan Das, Senior Applied AI Scientist, Manlike

Info: This session reframes "Loop Engineering" as applied cybernetics — the same observe→diff→act reconciliation pattern that powers Kubernetes controllers, now applied to AI agent loops. Attendees will learn how to map K8s design principles (declarative goals, idempotent actions, crash-only design, admission webhooks as HITL gates) directly onto agent infrastructure, and walk away with a concrete checklist for building production agent loops that self-heal, externalize state, and fail gracefully. We'll cover real failure modes — thrashing, stale reads, infinite reconciliation, conflicting controllers — and the fixes that transfer from a decade of K8s production learnings.


🕑: 11:45 AM - 12:30 PM
Automating Prompt Ops in a LLM World
Host: Afseen Syeda, Lead Prompt Engineer, Wisedocs

Info: Managing 10,000+ prompts across dozens of use cases is a challenging task when new LLMs get released weekly. This talk covers the infrastructure, automation and LLM based prompt that allows Wisedocs to automatically update prompts with new product releases, classify error categories and empower prompt engineers to focus on emerging, challenging problems, rather than maintanance.


🕑: 01:30 PM - 02:15 PM
From Explainable Evidence to Intelligence that Continually Learns
Host: Prashanth Rao, Founding AI Engineer & Researcher

Info: What if AI systems could learn continuously without repeatedly retraining the model underneath them? At HDC Labs, we are exploring hyperdimensional computing as a representation layer for online, associative learning. Neural models provide powerful learned representations and general capabilities; hyperdimensional representations provide a lightweight structure in which new observations, relationships, corrections, and outcomes can be incorporated as they arrive. Because learning can happen through simple composable operations over these representations, useful behavior can emerge from relatively few examples rather than large retraining datasets.

That same compositional structure also creates an opportunity for greater explainability. Instead of treating every update as an opaque change to model weights, hyperdimensional representations can preserve the primitives and associations that contributed to a result, making retrieval and reasoning easier to inspect. In this talk, we will


🕑: 11:45 AM - 12:30 PM
A Code-First BI Architecture Powered by AI
Host: Saeid Abolfazli, Global Head of Data Platform and AI

Info: When your BI team has 1000+ workbooks and a growing request backlog, the answer isn't more analysts — it's rethinking the architecture. In this talk we share how Rakuten Kobo's Data team replaced a centralized BI bottleneck with a code-first model where dashboards are version-controlled files authored end-to-end by AI. We cover the [long-term] strategy, tooling decision, the AI-guided authoring workflow built on Claude Code and Cloud, the real operating boundaries discovered through production-scale data. Attendees will leave with a concrete architecture pattern, an honest account of where the approach has limits, and a staged AI coding model they can adapt to any BI environment.


NOV 6

🕑: 10:55 AM - 11:40 AM
How to Not Be Wrong About AI
Host: Greg Wilson, Consultant, Third Bit

Info: Most organizations don't know how to assess the productivity of their software developers, which means that many claims about the impact of AI on productivity are vacuous. This talk analyzes some common mistakes, and describes a few things companies can do to figure out what is and isn't actually cost-effective.


🕑: 10:55 AM - 11:40 AM
TBD
Host: Mohammad Danesh, Head of Data and AI, Tangerine

Info: TBD


🕑: 11:45 AM - 12:30 PM
How to Measure Success for your AI Agents
Host: Manav Shah, Founding ML Engineer, raindrop.ai

Info: To turn the art of building AI agents into a science, you need a way to measure progress — and measuring an agent turns out to be far harder, and far more important, than measuring a model. This talk is about how to actually evaluate agents: what to measure, how the field gets it wrong, and why good measurement is what eventually lets agents improve themselves.

What we'll explore:

- What you're really testing when you evaluate an agent — the model vs. the harness around it
- The building blocks of an eval: traces, datasets, and scoring — and the dimensions most people miss
- Where today's benchmarks and LLM-as-judge break down, with real examples
- A practical playbook for measuring agent quality in development and in production
- How measuring success well opens the door to self-improving, self-healing agents


🕑: 11:45 AM - 12:30 PM
(PANEL) AI Sovereignty
Host: Diederik van Liere, Chief Technology Officer, Wealthsim
🕑: 11:45 AM - 12:30 PM
Responsible AI Beyond Accuracy: Fairness, Safety, and Sustainability
Host: Shaina Raza, Applied ML Scientist Responsible AI

Info: TBD


🕑: 01:30 PM - 02:00 PM
When Self-Improving Agents Learn the Wrong Lessons
Host: Abhimanyu Anand, Sr. Data Scientist, Elastic

Info: Most teams building self improving agents follow a similar trajectory: the agent runs, reflects, updates an artifact, and the metric rises. In production, the metric can still rise while a later incident traces back to a learned change that passed review but was evaluated against the wrong signals.

This talk examines different self learning systems that may change the agent context, the agent harness, or model weights. Each offers different gains and failure modes. Memory can absorb sycophancy and inflate its reward signal. Skill libraries can overfit to the last task or be tainted by one bad run. Harness changes can discard useful behaviour as new tools arrive. Weight updates persist longest and are the hardest to reverse.

I’ll cover the controls that make self iteration viable in production: bounded, reversible updates; provenance for every learned artifact; promotion gates; and evaluations that separate improvement from noise.


🕑: 01:30 PM - 02:00 PM
Grading the Agent: How We Built Evals for Disco's MCP Server
Host: Wendy Foster, Principal Data Scientist, Disco

Info: Disco's MCP server lets learning operators and administrators manage and generate insight about their programs through AI agents instead of dashboards, which means those agents need to be trustworthy, not just capable. This talk walks through how we built our eval process for MCP tools, from defining what 'correct' looks like for operator-facing actions to structuring a dataset that catches failures before operators do.


🕑: 02:05 PM - 02:35 PM
Recursive Self-Improvement
Host: Shashank Shekhar, Research Engineer, Google DeepMind

Info: Will provider later


🕑: 02:05 PM - 02:35 PM
Beyond Public Benchmarks: Building Enterprise AI Evaluations That Actually Pre
Host: Erin Li, Head of AI Research, CIBC

Info: Most AI benchmarks are optimized for public datasets and generic tasks, but enterprise teams operate in a very different reality: domain-specific workflows, regulated data, multimodal inputs, and business-critical quality thresholds. This talk presents the CIBC Benchmark, an enterprise evaluation framework designed to test AI models and workflows against real enterprise tasks, enterprise-relevant datasets, and standardized evaluation metrics.


Advertisement

Event venue & nearby stays

RBC WaterPark Place, 88 Queens Quay West, Toronto, Canada

Tickets CAD 0.00 to CAD 471.90
Concerts, fests, parties, meetups — all the happenings, one place.

Discover more events by tag

Ask AI if this event suits you

Keep the plans coming

More events in Toronto

Journey

Wed, 04 Nov at 07:30 pm Scotiabank Arena

Toronto is happening.

See everything else that’s on — concerts, markets, comedy and more.