The future of artificial intelligence looked bleak.
The field had spent the 1950s and 1960s promising machines that would translate languages and reason like people, but had failed to deliver any of it, and had watched governments on both sides of the Atlantic withdraw their funding through the 1970s in what the field itself would eventually name an “AI winter.”
What survived was expert systems, which worked by having engineers interview specialists and write down their knowledge as thousands of hand-written rules, and which were enjoying a brief commercial boom that would collapse within a decade. Neural networks, the idea of building something that learned from examples rather than being told the rules, had been effectively dead since the late 1960s, when an influential book demonstrated the limitations of the simple versions and the money went elsewhere. The people still working on them could be counted by hand, and most of them had been advised to work on something else.
By the 1990s, the words “artificial intelligence” had become slightly embarrassing to use. Researchers who wanted funding called what they did machine learning and borrowed their methods and vocabulary from statistics, and they talked about specific narrow tasks rather than about intelligence, which was what the previous generation had promised and failed to deliver.
None of this reached a boy in a town outside Montreal, for whom the words still meant whatever he had seen in video games.
Childhood
Hugo’s parents were both Québécois, descended from settlers who came over from France generations back. His mother stayed home through most of the years he was in school and later returned to the education system as an éducatrice spécialisée, working alongside children with learning disabilities and other difficulties, supporting them in the classroom. His father was a mécanicien d’entretien, a maintenance mechanic, the person a factory calls when a machine breaks or the line stops working. Nobody in the family had been to university.
Mathematics was Hugo’s favourite subject at school. He was curious about computers from a distance, since the family did not own one. He worked on convincing his father, insisting that they should buy a computer, until his father gave in and one came into the house. He then barely programmed on it at all.
His programming life began from a calculator. Graphing calculators had started appearing in Quebec high school mathematics classes, and teachers spent time showing students how to plot functions and solve equations with them. Buried in the manual was a programming feature, a simple language that let you write instructions the calculator would carry out in sequence. Around his fourth year of secondary school he decided to find out what he could build with it, so he built a map drawn out of text characters, with a figure standing on it who could be moved from one tile to the next, where arriving at certain tiles meant finding a treasure to pick up. He wrote a small puzzle story around the movement. It was the first thing he ever coded, it took him roughly a week, and he was proud enough of it to show his friends and several of his teachers.
There was, however, nowhere at school to take it further. He wanted to enroll at CEGEP in the computing program and pushed his parents hard on it. They pushed back, having been told by a teacher during a school visit that a student strong in the sciences would be bored inside a technical computing diploma, so he went into pure sciences instead. It annoyed him at the time. The computing electives on offer at CEGEP taught Word and Excel, and the only real programming course he took before university was a single class. But, his parents turned out to be right. The mathematics he took instead included content that was nearly impossible to learn later, and the programming was not.
Prior to CEGEP and during, Hugo had spent years watching American sitcoms to learn English. The Fresh Prince of Bel-Air was the first one he watched, and he could never say why afterwards, only that something in the way they spoke made it possible for him to understand. He followed one episode, then followed them without effort, and some time later he passed an episode of Home Improvement, which he had failed to follow before. He watched an episode and found he could understand all of it.
The show he loved most was Whose Line Is It Anyway?, a 1998 American improv programme airing on ABC in which performers were given a suggestion and had to build a scene out of it on the spot, and he watched it more or less alone, because it was an anglophone show and nobody he knew had heard of it. Then he mentioned it to a girl at CEGEP. She told him she loved it too and that she had never met anyone else in Quebec who did. On their first date they went to the cinema, got the showtime wrong, and missed the screening entirely, so she came back to his house (which was his parents’ place as he was still living with them), and they watched his recordings of the show instead. He ended up marrying her years later.
In his second year of CEGEP he entered a mathematics competition open to colleges across Quebec. He did not win, though. He finished somewhere around 20th or 30th, which was not high enough to mean anything, but high enough to clear the bar for a summer school at the Université du Québec à Montréal, where professors brought in from several universities taught the students who had placed well more advanced material than any of them had seen.
Before that summer began he had to choose a university. McGill and the Université de Montréal both accepted him, and McGill offered the better package, a scholarship renewable every year against a single year of funding at Montréal. Anyone asked which of the two carried more weight internationally would have said McGill without hesitation. However, he turned it down. He was not confident his English was strong enough to survive four years of coursework in it, and he was afraid he would struggle, so he enrolled at the Université de Montréal in computer science.
Then he went to summer school. One day, he sat down to lunch with a professor from the Université de Montréal, and she asked what he intended to study. He said computer science, and told her he was interested in artificial intelligence, a term he could not fully define, having met it mostly through video games and found it intriguing without knowing what it entailed. She told him there was a great deal of mathematics inside artificial intelligence, and that the joint mathematics and computer science degree would serve him better than the one he had signed up for.
He had weeks left before the term began. He changed his degree on the strength of one conversation across a cafeteria table, with a woman he had met that week and would never study under. It was an impulsive decision, and out of character, and he has never been able to explain why he made it at that time.
Université de Montréal
In his first year, a professor teaching one of the programming classes noticed his grades and told him about a federal program that funded undergraduates to spend a summer inside an academic lab. The pay was roughly $6,000 for three months, which was not a fortune but better than his job at McDonald’s in every possible way. He applied for the program, got accepted, and quit his job flipping burgers. The internship was in RALI, the natural language processing lab at the Université de Montréal, under Philippe Langlais, and it was his first exposure to a research problem.
However, it was not where he originally wanted to be. He had looked up which professors at the university worked deeply on AI, and one name came back in the search results. He emailed Yoshua Bengio to ask whether he was taking summer interns. Bengio replied that he did not have enough mathematics yet and that it was too early for him.
Some time later Hugo went to Bengio’s office to ask which classes he ought to take, and Bengio listed them. Almost all of them were statistics. Introduction to statistics first, then linear regression, then sampling, then a course on statistical estimation that spent the term on the properties of estimators, then statistical testing. Hugo walked out depressed, holding a list of the courses he would least have chosen for himself, because statistics was the part of mathematics that he enjoyed the least. The first course handed out formulas without explaining where they came from, which made the whole subject feel arbitrary and engineered rather than derived. He took them anyway. By the second course the probability arguments behind the formulas started showing up, and by the time he reached linear regression he was enjoying himself, having worked out that he was looking at a very simple form of machine learning and that you could say mathematically grounded things about how a predictive model behaves.
He came back the following summer, and Bengio took him under his wing. Another student arrived at the same time, a friend from the computer science program who had not taken mathematics, and who gave up after about three weeks and found an internship elsewhere. This was the summer Hugo first learned what a neural network was and what backpropagation did. He found it hard, but made it through, and he is certain he would not have without the dual degree he switched into.
The lab he had walked into was working against the consensus of the field. The methods that dominated machine learning then were kernel methods, methods associated not with non-convex optimization but convex optimization. This was considered a big breakthrough, being able to build a nonlinear model and frame it as a convex problem, because then you didn’t have to deal with difficult optimization problems. The methods also came with convenient mathematical properties that neural networks did not have, and a body of theory characterizing how well they would generalize. Their limitation was that they generalized locally. A kernel method handles a new photograph by leaning on the training photographs nearest to it, which works as long as the new thing sits close to something already in the collection. Show it a mango when it has only ever seen apples and oranges, and it will reach for whichever of those looks closest and call it that, with no way of registering that it is looking at a different fruit altogether.
Bengio believed that the biggest challenge in the field was finding methods that could generalize further from the training data than local kernel methods could manage, and that this was what the lab was trying to tackle from the beginning. His answer was the neural network, which learns a general account of the data rather than storing examples and comparing against them, and which might therefore say something sensible about a mango. Very few researchers anywhere were still working on them. They were regarded as a black box that took a kind of dark magic to make work, and success with them was thought to depend on unreliable insider knowledge passed hand to hand.
Bengio ran his lab as an open floor. If he had an idea he said it to everybody at once and hoped somebody would pick it up and run with it, which gave the room a strong sense that everyone was working on the same problem, and occasionally produced the discovery that two people had been chasing the same idea for a month. Hugo stayed there through a master’s, then a direct transfer into the doctorate on a scholarship, then the doctorate itself. He never seriously considered leaving. He had caught the bug, and there was too much left to learn.
The other thing he learned in that office had nothing to do with mathematics. He was sitting with Bengio one afternoon when the phone rang, and Bengio picked it up and began speaking French in a Parisian accent, which Hugo had never heard him do, having only ever known him with a Quebec accent. He was thrown by it. For a Quebecer, the French accent carries a whiff of pretension, and he sat there wondering who was on the other end of the line and why they warranted a different voice. He learned later that Bengio had been born in France and partly raised there before doing his schooling in Quebec, and that the switching was not performative but reflex, tuned to whoever had spoken last. In a three-way conversation with Nicolas Le Roux, a French classmate doing his PhD alongside him, Hugo would watch Bengio flip accents back and forth across a single conversation depending on which of them had opened his mouth. At the time it struck him mostly as evidence of how small his own world was. He had never been outside Canada, nor had ever been on an airplane. The first flight he ever took was to a machine learning conference, and he was nervous about it.
NIPS Conference
The conference was NIPS, around 2004, and he went with no paper and nothing to present, purely to read posters and find out what the field looked like.
Almost nothing there was about neural networks. He remembers something like four posters on the subject, off in a corner, and the rest of the hall given over to kernel methods, support vector machines, nonparametric Bayesian models, and a great volume of theoretical work proving generalization bounds. He could not read most of it. Bengio had taught his students neural networks and comparatively little of the learning theory everybody else in that room had grown up on, so Hugo spent the week walking past work he did not recognize, written in a language he had not learned, all of it more fashionable than his own.
“Am I studying the right thing?” he remembers thinking. “That doesn’t seem to be what other labs are studying.”
He was studying the right thing, and it would take the better part of a decade to prove that. In the meantime, the fact that so few people were doing it meant his expertise was worth far more when the field finally turned.
His work in those years went after a specific weakness in the kernel methods. The task was density estimation, modelling the distribution of some dataset of inputs, which at the time was still largely unsolved. The kernel approach placed all training examples and assumed a small Gaussian around every training example and assigned equal probabilities in every direction, quickly decaying as you got away from the training example, and taking the average of all these Gaussians, like a huge non-parametric Gaussian mixture. The problem was that a model given a pile of images had no way of working out which variations on them counted as real and which did not. A handwritten digit can be rotated, shifted in the frame, or drawn with a thicker stroke and remain the same digit, but the ways it can change are tiny in number next to the ways its pixels could be rearranged.
Bengio’s earlier answer was to work out those permitted variations separately around each example, using its nearest neighbours. Hugo’s idea was that the structure had regularity to it, since any digit can be rotated and any digit can be shifted, so a common structure ought to sit underneath all of it, and one neural network could produce those directions across the whole space rather than rediscovering them example by example. The paper, Non-Local Manifold Parzen Windows, came out at NIPS in 2005. Its central experiment taught the model about rotation using nine classes of digits, then handed it a digit 1, which it had never seen rotated, and asked it to turn 1. It could, even rotating in the opposite direction from anything in its training data, while the local method it was measured against could not. The paper was part of what Bengio’s lab was doing more generally, which was to take the popular kernel methods, find the problems they solved badly, and show that a neural network solved them better.
PhD
Nobody was training with GPUs in 2006, so the field was looking for algorithmic contributions that would help it train larger networks with less compute. Geoffrey Hinton had proposed one. Rather than training a deep network all at once, which was hard to do, you pre-trained each layer individually, one layer after another, using the learning algorithm of a restricted Boltzmann machine.
Bengio’s lab was exploring the idea in many different directions, and Hugo’s contribution was to suggest replacing the restricted Boltzmann machine with something simpler, for a reason he states without embarrassment, which is that he did not understand restricted Boltzmann machines very well at the time and was looking for a way around them. What came to mind was a single layer that takes an image, passes it forward into a hidden layer, passes it forward again into an output layer the same size as the input, compares the reconstruction against the original, and adjusts itself according to the difference. For instance, if you give a photograph of an apple and ask it to reproduce that apple, and it succeeds then it has learned something about how apples are put together.
He did not know at the time that the idea had been explored before, under names like “auto-associator,” or that “autoencoder” would become the term everyone used. It did not work as well as the restricted Boltzmann machine, but it worked considerably better than doing no pre-training at all, and it was one of the first things he felt he had contributed from his own ideas that seemed to have legs. It went into the paper, Greedy Layer-Wise Training of Deep Networks, published at NIPS in 2006.
His version had a problem. If the hidden layer was as big as the input layer, nothing stopped the network from learning a representation that was simply a copy of the input, taking each individual pixel into its own hidden unit, copying it out again, and solving the problem trivially. Pascal Vincent was a junior professor in the department who had done his own PhD with Bengio, and Hugo happened to go into his office one day to talk about the work he was doing. Vincent suggested that a way of avoiding the copying solution would be to add noise to the input, so that the problem the network had to solve was not to reconstruct the original input but to uncover some of the noise that had been added to it. Hugo thought immediately that it made intuitive sense, and that a network doing that would probably need to learn something about the regularities in the input distribution and about what makes different pixels depend on one another.
The first time they tried it, it just worked. What they would do at the time was train on handwritten digits and then look at the filters learned by the hidden units, and what you were hoping to see, which is something you would see with restricted Boltzmann machines, were stroke detectors, filters that looked like a pen stroke with empty background around it. They were not getting that with plain autoencoders, which they thought was because the model was too tempted to copy each individual pixel. As soon as they added a bit of noise, the stroke detectors started to appear. The work went to ICML in 2008 as Extracting and Composing Robust Features with Denoising Autoencoders. Hugo is careful to say that at the start neither of them really knew why it worked, that the more formal account of how the method relates to the local structure of the data distribution was shown afterwards, mostly by other people in Bengio’s lab.
The other paper from those years is the one he regrets, though not for its contents. Hugo, Bengio, and a fellow student had built a system that could produce a classifier for a task it had never seen, working from a description of the task rather than from labelled examples of it, and they called this Zero-data learning. Almost immediately afterwards somebody else published nearly the same idea under the name “zero-shot learning,” which is the name the entire field uses today. For years afterwards people kept asking Hugo what zero-data learning could possibly mean, since learning from no data is a contradiction.
NIPS rejected the zero data-learning paper. Hugo sent it instead to AAAI, whose deadline fell a week before ICML’s, which suited him because it cleared one paper out of the way before he pushed the denoising autoencoder to ICML. The venue change also let him write a sentence he could not have written for a machine learning audience. In the mid-2000s the term artificial intelligence was very nearly unsayable in the rooms Hugo worked in, contaminated by expert systems and the failures of the previous generation, and the machine learning community had responded by borrowing heavily from statistics, narrowing its ambitions, and talking about tasks rather than about intelligence. AAAI had the words in its name. So in the conclusion, he and his co-authors wrote about extending machine learning towards AI, and described a system that would understand instructions specifying a task and then perform that task with no additional supervised learning, provided it had already been trained on related tasks described in the same language.
He was describing prompting, but he would not see it work for another fifteen years.
Toronto
He reached the end of his PhD without a plan.
He knew he liked research and wanted to stay in Quebec, and staying in Quebec meant academia, because there was no industrial lab in the country doing the work he cared about. No Google AI lab, no DeepMind, nothing. So he told Bengio he would do a postdoc and asked whether he should write to Hinton. The two were not collaborators but they spoke regularly, and the arrangement came together almost immediately, much faster than these things normally do, which Hugo puts down to a phone call in which Bengio presumably said something generous. He applied for a fellowship, won it, and moved to Toronto.
He was a nervous traveller by his own description, and living in English made him nervous even inside his own country. He had a daughter by then, born during the PhD, and a second would arrive while they were in Toronto, and he did not want to be far from his parents, who would be needed. The United States was the obvious move for someone in his position, but he did not seriously consider it for those reasons.
Hinton ran a lab that was the inverse of Bengio’s. Where Bengio broadcast ideas into a room and let them find an owner, Hinton gave each student and postdoc a project that was theirs from the first day and understood by everyone to be theirs, which Hugo found clarifying. His fellowship was generous enough that his salary was not coming out of Hinton’s grants, which bought him room to work on things that were not Hinton’s projects. He used it to join Rich Zemel’s reading group and finally learn the probabilistic graphical models that Bengio’s lab had never taught him.
The problem he picked was one that had frustrated him. There was no reliable way of evaluating an unsupervised model. A lot of the methods that existed at the time were looking at the filters a model had learned and forming an indirect judgment about the features. An apple is a useful way to see what was missing. A model that has learned what apples look like ought to rate a new photograph of an apple as ordinary and a photograph of scrambled pixels as nearly impossible, and that rating is the number nobody could get out of these models directly. Ruslan Salakhutdinov had started scoring the quality of these models by asking how much probability they assigned to real images they had never seen, which was the right idea, except that computing it involved an approximation, and the approximation was what frustrated Hugo. Working with Iain Murray, another postdoc in Hinton’s lab, he found a way to build a model that produced the number exactly. Rather than having the model describe a whole image at once, they had it describe one pixel at a time, each pixel predicted from the pixels before it, which meant the total could simply be multiplied out. The result was NADE, The Neural Autoregressive Distribution Estimator. It was the first paper he published with no supervisor on it, made by two postdocs in a room, and the impact it had was theirs.
He also spent time on an idea of Hinton’s about foveation, building a model that chose where to look inside an image, took a glimpse at that spot, accumulated what it had seen, then used the accumulated glimpses to decide what the image contained. Everything in that lab was an attempt to train bigger networks on more data through cleverness rather than force.
Hugo had left Toronto by the time AlexNet won the ImageNet competition, realizing how important it was to the field despite never sharing a conviction on it. Ilya Sutskever had become persuaded that the answer was to train a much larger model on a much larger collection of images and solve whatever engineering problems stood in the way, which meant abandoning the belief that a deep network had to be assembled layer by layer before it could be trained at all. Hugo had spent years helping to build that belief. The layer-by-layer methods turned out to be unnecessary. What the problem had actually needed was a better training procedure and a graphics card fast enough to run it.
YouTube Channel
Hugo returned to Montreal and took a faculty position at the Université de Sherbrooke, and discovered on arrival that he had never taught anything to anybody.
This is ordinary for new professors and nobody warns them about it. The first thing he learned, which he says almost everyone learns in their first semester, is that what is obvious to you is not obvious to the students, and that the correction feels like dumbing the material down when it is really just putting it at a reasonable pace. He was lucky in what he was asked to teach. They gave him introduction to artificial intelligence, then let him build courses on neural networks and machine learning from scratch, and never made him teach introductory programming.
Open online courses were becoming a phenomenon at the same moment, and Hugo looked at them and recognized himself as a student. He would have loved it, being able to watch a lecture at his own speed, stop it, watch a section again, sit with something difficult until it opened, then bring a question to a professor already knowing where his confusion was. So he inverted his classes. The lectures went onto video and the students watched them at home, and the class hours were given over to exercises, with him in the room while they worked, which is the arrangement known as the flipped classroom and which Sherbrooke’s department let him try. The results were at least as good as what he had been getting the conventional way.
The videos went onto YouTube, where they were watched by a great many people who were not his students, and this had an effect he had not anticipated. People began recognizing him at conferences and asking to take photographs with him. He had become a known name in deep learning partly through his papers and partly through a course he had recorded alone in front of a laptop.
Recording was strange at first, sitting in a room talking to nobody, and then he began to enjoy it, including the editing. Hugo walked into every lecture that term more prepared than he had ever been in his life, because a stumble would be visible forever and painful to edit and might mean shooting the whole thing again. He says he was never especially good at teaching, that he put himself in conditions where he was energized to do it, and that innovating on the format made it feel a little like research.
Hugo does not teach any more. He still answers questions under the videos, which scratches the itch, and he still gets comments telling him his French accent is strange, which he says he has learned to go beyond at this point.
Whetlab
During the Toronto postdoc he met Ryan Adams and Jasper Snoek, and together they set out to automate one of the most tedious jobs in the field. Every machine learning model has hyperparameters, settings like the learning rate and the number of units, which have to be chosen before training starts. At the time, people saw the training of a neural network as a black box and the researcher picked those values by hand. Their method used Bayesian optimization to choose them automatically instead, and the experiment that struck people was one in which it tuned the hyperparameters of a deep neural network better than the original settings the researcher had arrived at by hand.
The paper, Practical Bayesian Optimization of Machine Learning Algorithms, was finished after Hugo had moved to Sherbrooke, which makes it technically his first Sherbrooke paper despite having begun in Toronto, and it went to NIPS and very nearly did not get in. The reviews ran from lukewarm to negative. One said the work was not particularly creative, on the grounds that the underlying technique already existed and the authors had merely pointed it at a new target. The area chair disagreed, championed the paper against its own reviews, and pushed it through.
The machine learning community then decided it was one of the most useful things published that year. It became one of Hugo’s most cited papers, and at that NIPS it finished as the second most cited paper of the entire conference. The most cited paper that year was the ImageNet paper out of Hinton’s lab.
Interest followed, particularly from the large technology companies, which wanted anyone who could do this kind of work. Ryan Adams looked at the situation and concluded they should start a company. He convinced Hugo and Jasper Snoek, recruited Kevin Swersky, a Toronto PhD student working on the same problem, and brought in Alex Wiltschko, a neuroscientist who had met Adams and could actually build things.
Hugo was a part-time founder. He was still a professor with a lab and a teaching load, and he says his contributions were minor and the credit belongs to the others, who pushed hard enough to get a real product working, a service customers could query to receive tuning recommendations. He gave some input on product design and found the company its first client, a deal that never closed, because the company was acquired before the contract was signed.
The acquisition was of its time. There was so little deep learning talent in the world that the large companies were buying teams as readily as technology, and Twitter, which had people inside it arguing that it needed a deep learning laboratory of its own the way Google and Facebook had theirs, bought them. Hugo believes the technology was used, and also that the purchase was partly a signal, a way of announcing that Twitter took this seriously, helped along by the fact that several of them came out of Toronto.
Jack Dorsey returned as CEO in the middle of the negotiation. The founders panicked for a while about whether a change at the top would kill the deal. It did not, and Hugo met Dorsey a handful of times afterwards.
He took the offer for two reasons. The first was financial. The second had been bothering him for years. He had never worked in industry, never done an industrial internship, never built anything that had to be deployed, and he was training students who almost all went into industry, in a field that exists to solve concrete problems rather than abstract ones. He felt strange about not having that experience himself. Sherbrooke gave him a leave of absence, let him keep his affiliation and finish supervising his students, and he moved his family to Boston.
Game of Thrones
The research he did at Twitter came out of watching engineers do something tedious.
Engineers designed classifiers of tweets by hand, and they had to design a lot of them, because that was one way to track specific types of content on the platform or surface certain signals. Somebody would ask whether they could get a classifier about Game of Thrones, because the product team wanted to surface tweets about it when the latest episode had just aired. Hugo looked at that and saw a value in designing methods that could produce classifiers automatically from very few examples of the kind of tweets you wanted.
So he started on few-shot learning, which means building systems that can generalize from a handful of examples, using a family of methods called meta-learning, where you design learning algorithms that discover other learning algorithms. It was related to the hyperparameter work in Toronto, since both were attempts to automate part of a machine learning scientist’s job, except that this time the target was not the settings but the learning rules themselves.
The paper he published from Twitter, Optimization as a Model for Few-Shot Learning, contained no tweets at all, and tweets were its entire motivation. He is prouder of it than of almost anything else he has written, partly because of where it took the field, and partly because, years later, the Google Vice-President Blaise Agüera y Arcas gave a keynote at NeurIPS 2019 and said the work had inspired his own.
His second assignment at Twitter was to build a research culture inside a company that did not have one, and this is the part that failed. Twitter had not turned a profit in years and there was open speculation in the market about whether it ever would, and eventually the decision came down that there would be no separate research laboratory, that it would be folded into the product organization. Hugo realized that whatever had been brought in to provide was no longer wanted.
The timing was almost absurd. In the same period, Samy Bengio told him that Google wanted to open a Brain laboratory in Montreal, the first Google Brain team anywhere outside the United States, and was looking for someone with the experience and the appetite to build a team from nothing.
Google Brain
Google chose Montreal because of Montreal, which is to say because of Bengio and the ecosystem that had accumulated around him, and because reinforcement learning was becoming commercially important at exactly the moment when Doina Precup and Joelle Pineau were both working in the city. Hugo’s reason was that his family would be home.
There was already a Google office, so the office setup had already existed and he would not need to spend a year negotiating for desks. What he was given from the get-go was headcount. His first recruits were Marc Bellemare, Nicolas Le Roux, and Danny Tarlow, and two of the projects he is proudest of from that decade came from them.
Bellemare took on Loon, a Google moonshot that flew balloons through the stratosphere to deliver internet connectivity, and replaced the balloons’ navigation system with one that had taught itself to steer. Learning by trial and error was widely regarded then as a technique that worked beautifully in simulation and fell apart on contact with the physical world, and here was a fleet of real balloons crossing the globe under a learned controller. The work went into Nature.
Tarlow wanted to apply deep learning to code, which in the late 2010s was a niche interest almost nobody was pursuing. The first thing his group built was a system that suggested fixes when a build broke, trained on Google’s own record of what its engineers had actually done in that situation over the years, on the reasoning that a task performed constantly and captured in data is a task a model can learn. Tarlow is still at Google and is now central to Gemini’s coding abilities, and Hugo says with some satisfaction that his team was years ahead on what turned out to be the most popular use of AI in the world, today.
His own role had changed shape, and he has two ways of describing it. When he was a PhD student writing papers he owned, he was an artist making albums, and he had since become a producer, making other artists’ records. The other description is gardening. The researchers around him had doctorates and reputations and did not need direction, and telling them what to work on was neither possible nor desirable, so the job was to prepare the conditions in which good research grows, meaning the physical space, the ease of collaboration, and the chance encounters between people from different fields.
Through all of it he kept one day a week at Mila, where he held an adjunct professorship and supervised students, some jointly with professors there. He had been trying to knit the Montreal ecosystem together since he was a graduate student, when he and Doina Precup started alternating their two labs’ seminars between the Université de Montréal and McGill, for no reason beyond its being absurd that two groups working on the same problems in the same city should sit separately. Mila became a national institute in 2017, alongside Vector and Amii. Hugo’s view of his own position was that it mattered not only for Google to be in Montreal, but for there to be some Montreal inside Google.
Conferences
Somewhere in this period he became one of the people who ran the field’s institutions, and then became the person trying to change them.
He organized ICLR for several years and sat on its board, sat on ICML’s board, and eventually served as senior program chair of NeurIPS, the largest and most prestigious conference in machine learning, by then an event with thousands of submissions and an army of reviewers who had to be recruited, evaluated, and replaced. With the conference co-program chairs, he built the reviewer pool by pulling in everyone who had been an author on a recent NeurIPS paper, scored the reviews people wrote so the worst ones would not be invited back, and capped each reviewer at six papers.
After years of these conferences, Hugo had concluded that a conference is asked to do two jobs at once and does both badly. The first is to decide whether a paper’s claims are actually supported by its evidence, which is a question with an answer. The second is to decide which of the papers that clear that bar are exciting enough to feature, which is a matter of taste, and entirely legitimate for an event that exists partly to put on a show. Stapling the two together produces most of the randomness in accept and reject decisions, and it produces the particular misery of an author who cannot work out why a reviewer is unenthusiastic about work nobody has claimed is wrong.
So he launched a venue that would do only the first job. Transactions on Machine Learning Research launched in 2022, with Hugo and Kyunghyun Cho and Raia Hadsell among its editors in chief and Fabian Pedregosa as its founding managing editor, reviewing in the open against a single criterion, which is whether the claims match the evidence. Excitement is handled separately, through certifications attached to papers that have already been accepted, and the ambition is that conferences will eventually award those certifications themselves, so that a paper can be checked once and then featured by whichever communities find it interesting. The first venues to try it were small, AutoML and CoLLAs. And now, ICLR, ICML, and NeurIPS allow for some TMLR papers to be presented at these conferences, and Hugo feels quite proud of having made that possible. Submissions are accepted at any time, which dissolves the deadline cycle that has every researcher in the field cramming toward the same four dates a year and running out of compute in the same week. It reached roughly a thousand submissions a year, and a year later, Hugo said it no longer felt like an experiment.
Mila
What ended his time at Google was publishing.
He had joined on the understanding that he would be free to publish, and for years that was uncontroversial, because Google was widely regarded as the place to do AI and had concluded it gained more than it lost by putting its research into the world. Then OpenAI released ChatGPT and demonstrated that somebody else might reach very powerful systems first, and Google grew steadily more careful about what it was willing to say in public. Hugo does not dispute the company’s right to make that call. It is a company, and it is up to a company to decide. It was simply no longer the job he had taken. He went part time so he could spend his own hours at Mila, not as a Google person but as a passion project on his own dime.
Two other things were happening at once. The relationship between Canada and the United States had soured, and he found it less appealing to be contributing to an American company, though he is careful to say Google always felt more international than American and that this alone would not have moved him. And Yoshua Bengio had decided to hand off the scientific direction of Mila so he could put his energy into LawZero, his AI safety project.
The professors at Mila approached him, and Hugo said yes to join Mila as the successor of Bengio in the role of the Scientific Director.
The role is the one he has been doing for a decade, which is gardening, except the garden is now an institute of more than 1,500 researchers, including students, with a community of associate professors who work at the intersection of AI and other disciplines and whom he treats as his own continuing education. When people ask whether it is intimidating to take Yoshua’s chair, he says it is an honour, that the next chapter is a different one, and that what he brings to it is industrial.
What he took from Google was not a technique but a posture, the ambition, the insistence that a discovery is supposed to become something real, and the observation that Canada has been a world leader in AI research and considerably worse at putting any of it to use. He wants Mila’s students to graduate feeling capable of starting companies. He wants to narrow the distance between what the research can do and what anyone actually uses, and he wants it narrowed in a way that is recognizably Quebec’s rather than Silicon Valley’s.
Vigilant
At home there are four daughters, and what they hear at school is that AI is something they are not supposed to touch. He thinks that is a conversation Quebec is going to lose. The rule should be about cheating, which is a rule that already exists and always has, and which does not change because the tool changed. So he keeps it unforbidden in his own house, telling his children that a question they are wondering about is one they could put to Gemini, pulling out his phone in social situations to see whether a machine provides a useful suggestion, encouraging them to use it when they are stuck on a concept and want to understand it rather than when they want to reach the answer quickly.
The word he uses when people ask whether he is worried is “vigilant,” and he says it is deliberate and that not enough people use it. The opportunities are real, the risks are real, and which of them you end up with depends on how the technology is adopted, so the right posture is constant attention. He thinks the field has done better here than social media did, in that the scrutiny arrived early rather than a decade late, and he considers that healthy.
The idea he keeps returning to lately comes from François Jacob, who divided scientific work into day science and night science. Day science is coding up the idea, running the experiments, and analyzing the results. Night science is deciding which problems are worth years of a life, forming hypotheses, choosing where to point everything.
Hugo has watched coding agents tear through day science at a speed that has unsettled his own students, who are asking, with some justification, whether what they are doing is research at all. He is comfortable with that part being automated, because the middle of a project was never the fun part. The joy is at the beginning, when the idea arrives, and at the end, when it turns out to have been a good one.
Night science is different, and his objection to automating it has nothing to do with whether a machine could. It is about legitimacy, about who is entitled to decide where a society spends its time, money and attention, and about wanting a person accountable for choices that large.
He thinks the field does not do enough of it. He includes himself, saying he has mostly done night science when the job demanded it, that he could be more ambitious about it, and that it is the part of the work he finds most exciting anyway.
He talks now about wanting Mila to be where the next real breakthrough happens, and he is honest that the past decade has mostly been the scaling of ideas worked out while he was a graduate student. He cannot make that breakthrough arrive alone.
What he can do is keep the room open, the same way it was open for a student who was turned away by Bengio for not having enough mathematics and came back the next year.



