Skip to content

Résumé

Sourena Khanzadeh

AI Research Engineer · Toronto, Ontario, Canada. AI Research Engineer focused on reliable agentic systems, agent evaluation, and auditable AI workflows.

Summary

AI Research Engineer with a Ph.D. in Computer Science and 3+ years of experience building and evaluating LLM agents, deep-learning pipelines, and search-based AI systems across industry, government, and academia. Builds agent harnesses, retrieval and tool-use workflows, evaluation infrastructure, and reproducible ML prototypes; published at AAAI, at Canadian AI, and in peer-reviewed journals.

Open to research collaborations and selective advisory work in agent evaluation and reliability.

Experience

AI Engineer, Agentic Systems

Flybits · Full-time · Toronto

Agentic systems that plan, call tools across internal and partner APIs, and complete multi-step tasks with every decision logged for end-to-end auditability.

  • Retrieval over structured user-context graphs, pairing learned models with symbolic constraints so responses stay grounded in real customer state rather than model priors.
  • The evaluation loop — task completion, groundedness, faithfulness, recovery from tool failure — and using it to tune model routing across capability, latency, and cost.

AI Research Engineer, sAIpien

MIT Media Lab · Part-time research affiliation

Perspective-aware agents, privacy-preserving long-term memory, and benchmarks for transparency and human oversight.

  • Prototyping perspective-aware agents and privacy-preserving long-term memory through adaptive knowledge graphs (Chronicles).
  • HCI² benchmarks for transparency and human oversight.

Postdoctoral Fellow / AI Research Engineer

Toronto Metropolitan University · Part-time academic appointment · Toronto

Reproducible evaluation harnesses for LLM-agent reasoning, orchestration, context management, and reliability.

  • Evaluation harnesses in Python and PyTorch — trajectory logging, counterfactual replay, automated scoring.
  • Taking research questions to working prototypes with experimental protocols, metrics, and publications, shipping code other researchers can run and reproduce.

Machine Learning Engineer (Graduate Research)

Toronto Metropolitan University · Toronto

AI prototypes spanning multi-agent systems, heuristic search and planning, deep learning, reinforcement learning, and GAN-based synthetic-data generation.

  • Training and inference pipelines with PyTorch, Hugging Face, scikit-learn, and XGBoost; experiments containerized with Docker and checked through GitHub Actions CI.
  • Controlled experiments with ablations and failure-mode profiling, published at AAAI, Canadian AI, and in peer-reviewed journals.

ML Engineer / Research Scientist Intern

National Research Council Canada · Toronto

Knowledge-informed machine-learning models for anomaly detection over large-scale, severely imbalanced telemetry data.

  • Domain-derived features that reduced dependence on large labelled datasets.
  • Reproducible, peer-reviewed experiments across an interdisciplinary team.

Selected projects

Preprint + public research artifact

An agent's written reasoning is not evidence that the reasoning caused the answer.

23 of 30audited trajectories showed a faithfulness violation under the counterfactual-intervention protocol. This is an audit finding about the agents under test, not a failure rate of the tool. 23/30 is 76.7%.

Prototype architecture and case study

A single LLM asked to build software stops at the first thing it gets wrong.

Self-repairexecution feedback drives automatic repair — demonstrated end to end, not benchmarked. No benchmark, baseline, or success rate is reported in the paper, so no comparative claim is made here.

Peer-reviewed research

When a heuristic gives a search no gradient to follow, the search stalls in a region it cannot see its way out of.

70.3%of unbounded-region tasks, where the restarting-random-walk planner was faster than or uniquely successful against breadth-first search. The 70.3% is measured on unbounded-region tasks specifically, not on planning benchmarks as a whole. "Uniquely successful" means the comparator did not solve the task at all.

Peer-reviewed research

A 210-image, severely imbalanced dataset is too small to train a classifier that generalizes.

91.5%downstream classification accuracy — 4 points over duplication-based oversampling, 5.5 over none. This is downstream classifier accuracy on the microplastics dataset, not a measure of the generated images themselves.

Technical skills

Languages
Python · C/C++ · TypeScript · Java
Machine learning
PyTorch · Hugging Face · scikit-learn · XGBoost · Deep learning · Reinforcement learning · Fine-tuning · GANs · Synthetic data
LLM & agent systems
LLM agents · Tool and function calling · Multi-agent orchestration · Retrieval-augmented generation · Context and memory management
Evaluation
Agent harnesses · Trajectory logging · Counterfactual evaluation · Automated scoring · Experiment design · Observability · Guardrails · Failure analysis
Infrastructure & software
Docker · Kubernetes · AWS · PostgreSQL · Redis · Git · GitHub Actions · APIs · Testing · Reproducible pipelines

Education

Ph.D. in Computer Science

Toronto Metropolitan University

Agentic AI workflows for reliable automated reasoning — LLM-agent orchestration, reasoning loops, tool use, and evaluation of agent faithfulness. GPA: A+. Coursework: Heuristic Search, Deep Learning, Directed Intelligent Robotic Systems.

Bachelor of Computer Science

Toronto Metropolitan University

A+ in every AI course: Machine Learning, Artificial Intelligence, Reinforcement Learning, and Computer Vision.

The complete record — peer-reviewed papers and preprints, labelled separately, with citations — is on the research page. The Google Scholar profile is authoritative for citation counts.