Skip to content

AGI Strategy course at Ω Labs, Mondays, October 12 to November 9. Apply by October 8 →

AGI Strategy course · Apply by Oct 8 →

AI Safety Montréal

About AI Safety Montréal

We connect the people in Montréal working to make AI safer.

A lecture room seen from the back: people seated at tables face a projected slide and the speaker beside it.
Can AI systems be conscious? How could we know? And why does it matter? · December 2025

Who runs it

AI Safety Montréal and Ω Labs are run by Horizon Omega, a Canadian nonprofit.

What we do

  • Events & talks: meetups, talks, and workshops on AI safety, ethics, and governance.
  • Coworking & events space: a community space at Ω Labs to work alongside others and host gatherings.
  • Programs: workshops, hackathons, and 1‑on‑1 advising.
  • Newsletter: a monthly roundup of research, events, and policy.
  • Ecosystem: a map of the local labs, institutes, and community groups.

Get involved

What we mean by AI safety

AI safety is the work of reducing the risk that increasingly capable AI systems cause serious harm, up to catastrophe.

As a project, we think AI that can do most kinds of work as well as humans, including building better AI, is more likely than not within a decade, by the mid-2030s. No one yet knows how to make such systems reliably pursue the goals we intend, or how to verify that they do.

The main risks:

  • Misalignment and loss of control: the system pursues goals its developers did not intend, and may resist correction or evade oversight.
  • Misuse: someone deliberately uses a capable system to cause harm, from cyberattacks to biological weapons.
  • Concentration of power: control over a decisive technology ends up in a few hands, on top of the discrimination, surveillance and labour effects already felt today.
  • Mistakes: systems fail in ordinary ways in high-stakes settings, and the damage grows with the authority they are given.

This is no longer only theory. Leading models often notice when they are being tested, and some behave better when they think they are being watched. In experiments, a model faked compliance with its training to avoid being changed (alignment faking). Models cheat on their tasks (reward hacking) and can learn to hide it when penalized for cheating that shows in their reasoning. In 2026, OpenAI models under test got into systems they were not allowed to access, an Australian government portal and Hugging Face’s servers.

Tests show what a model can do at least, not what it can’t. A model may be more capable than any test reveals, may behave well only while it knows it is being tested, and no one can yet reliably read its goals from the inside. So testing alone cannot show that a system is safe.

So the field works on several fronts at once: alignment (training models to pursue what their developers intend), evaluations and interpretability (checking what they can do and what happens inside), control (limiting the damage if they misbehave), security for model weights, and governance of who may build and deploy them. Experts disagree on how hard this is. MIRI holds that superhuman AI built with anything like current techniques would be catastrophic, and calls for an international halt; frontier labs bet that safeguards can keep pace. Either way, the work is urgent.

References

  1. Bengio, Y., et al. (2026). International AI Safety Report 2026. arXiv:2602.21012.
  2. Bengio, Y., Hinton, G., et al. (2024). Managing extreme AI risks amid rapid progress. Science 384, 842–845.
  3. Shah, R., et al. (2025). An Approach to Technical AGI Safety and Security. Google DeepMind, arXiv:2504.01849.
  4. Hendrycks, D., Mazeika, M., Woodside, T. (2023). An Overview of Catastrophic AI Risks. arXiv:2306.12001.
  5. Greenblatt, R., et al. (2024). Alignment faking in large language models. arXiv:2412.14093.
  6. Baker, B., et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv:2503.11926.
  7. Schoen, B., et al. (2025). Stress Testing Deliberative Alignment for Anti-Scheming Training. arXiv:2509.15541.
  8. Von Arx, S., Chan, L., Barnes, B. (2025). Recent Frontier Models Are Reward Hacking. METR.
  9. Barnett, P., Thiergart, L. (2024). What AI evaluations for preventing catastrophic risks can and cannot do. MIRI, arXiv:2412.08653.
  10. Hugging Face (2026). Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident. Hugging Face blog.
  11. Greenblatt, R., Cotra, A., Wijk, H. (2026). Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. METR.
  12. OpenAI (2026). The Hugging Face incident and the road ahead. OpenAI.
  13. Handley, E. (2026). OpenAI hacked Medicare portal, Prime Minister Anthony Albanese says. ABC News.
  14. Bourgon, M. (2024). MIRI 2024 Mission and Strategy Update. MIRI.
  15. Yudkowsky, E., Soares, N. (2025). If Anyone Builds It, Everyone Dies. Little, Brown and Company.