InfluencesPeers & partnersSuccessors
Hover for the evidence (or the center for a bio) · click to travel · drag a bubble to tug it · scroll to zoom
Curator
Drawing the map…
Yoshua Bengio: influences, peers and legacy
Every connection with its receipt
Yoshua Bengio's map is a lineage of connectionism: the 1980s backpropagation revival, the obscure recurrent-network papers that framed the vanishing-gradient problem he later formalized, and the Montreal ecosystem he built at MILA. It reaches beyond computing into cognitive psychology and causal philosophy, and ends with the attention-based architectures his lab's work directly seeded.
The Kynda mix for Yoshua Bengio
- Key Influence · Learning representations by back-propagating errors by David Rumelhart, Geoffrey Hinton, Ronald Williams (1986). The paper that pulled Bengio into neural networks as a McGill graduate student in the mid-1980s, when symbolic AI dominated and connectionism was fringe. Backpropagation gave him the gradient-based training method underlying nearly everything he later built, from neural language models to attention. He would share the 2018 Turing Award with Hinton and LeCun precisely for extending this line of work into deep architectures.
- Influencia Obscura · Finding Structure in Time by Jeffrey L. Elman (1990). Elman's simple recurrent network, built to show that grammar-like structure could emerge from sequence prediction, is the under-cited ancestor of Bengio's recurrent work. His 1994 paper on the difficulty of learning long-term dependencies is essentially a mathematical diagnosis of why Elman-style networks fail on long sequences — the problem whose eventual solutions became LSTMs, gating, and attention.
- Local Roots · Spoken Dialogues with Computers by Renato De Mori (1998). De Mori ran the speech recognition group at McGill where Bengio worked in the late 1980s, and speech — with its sequential, noisy, variable-length data — was the problem domain that pushed him toward recurrent models and hybrid neural/HMM systems. Montreal's concentration of speech and pattern recognition research in that era shaped the problems Bengio chose long before MILA existed.
- Beyond the Medium · Thinking, Fast and Slow by Daniel Kahneman (2011). Bengio built an entire research program around Kahneman's System 1 / System 2 distinction, arguing in his widely circulated NeurIPS 2019 keynote that deep learning had mastered fast intuitive perception but not slow deliberate reasoning. The framing drove his subsequent work on consciousness priors, attention as a bottleneck for conscious processing, and out-of-distribution generalization.
- Peer · Deep Learning in Neural Networks: An Overview by Jürgen Schmidhuber (2015). Schmidhuber's Swiss lab pursued recurrent networks through the same wilderness years as Bengio's Montreal group, and his LSTM work with Hochreiter solved the long-term dependency problem Bengio had diagnosed in 1994. His public disputes over credit with the Turing trio made the rivalry famous, but the parallel tracks — gated recurrence versus Montreal's attention and generative models — defined the field together.
- From the Canon · A Neural Probabilistic Language Model by Yoshua Bengio (2003). The paper that replaced discrete n-gram counting with learned distributed word vectors, inventing word embeddings years before word2vec made them ubiquitous. It is the clearest single statement of Bengio's core conviction — that meaning should be learned as continuous representation rather than engineered as symbol — and it is the origin point of the neural approach that now underlies every large language model.
- Key Collaborator · Explaining and Harnessing Adversarial Examples by Ian Goodfellow, Jonathon Shlens, Christian Szegedy (2015). Goodfellow was Bengio's doctoral student in Montreal, where he conceived generative adversarial networks — the idea famously sketched after a bar argument with labmates and published with Bengio as co-author in 2014. This companion paper on adversarial fragility shows Goodfellow's own signature move: probing where learned representations catastrophically break, a concern Bengio's lab made central.
- Legacy · Attention Is All You Need by Ashish Vaswani, Noam Shazeer, Niki Parmar et al. (2017). The Transformer paper takes the attention mechanism Bahdanau, Cho and Bengio introduced in Montreal and discards recurrence entirely, keeping only the soft alignment. It cites that 2015 work as the foundation it builds on, and every large language model since descends from this pairing — making Bengio's lab the upstream source of the architecture that now defines the field.
What influenced Yoshua Bengio
- Learning representations by back-propagating errors by David Rumelhart, Geoffrey Hinton, Ronald Williams (1986). Learning representations by back-propagating errors (David Rumelhart, Geoffrey Hinton, Ronald Williams) — titan for Yoshua Bengio “Hinton received the 2018 Turing Award, together with Yoshua Bengio and Yann LeCun, for their work on deep learning.” (en.wikipedia.org)
- John Hopfield. Deep Learning (Ian Goodfellow, Yoshua Bengio and Aaron Courville) — legacy for John Hopfield (via Deep Learning) “In 2025, Bengio was awarded the Queen Elizabeth Prize for Engineering jointly with Bill Dally, Hinton, John Hopfield, Yann LeCun, Huang and Fei-Fei Li.” (en.wikipedia.org)
- Parallel Distributed Processing by David Rumelhart and James McClelland (1986). Parallel Distributed Processing (David Rumelhart and James McClelland) — titan for Yoshua Bengio (wikidata.org)
- Thinking, Fast and Slow by Daniel Kahneman (2011). Thinking, Fast and Slow (Daniel Kahneman) — culture for Yoshua Bengio (openlibrary.org)
- Untersuchungen zu dynamischen neuronalen Netzen by Sepp Hochreiter (1991). Untersuchungen zu dynamischen neuronalen Netzen (Sepp Hochreiter) — ghost for Yoshua Bengio (wikidata.org)
- A fast learning algorithm for deep belief nets by Geoffrey Hinton, Simon Osindero, Yee-Whye Teh (2006). A fast learning algorithm for deep belief nets (Geoffrey Hinton, Simon Osindero, Yee-Whye Teh) — titan for Yoshua Bengio (wikidata.org)
- Gödel, Escher, Bach: An Eternal Golden Braid by Douglas Hofstadter (1979). Gödel, Escher, Bach: An Eternal Golden Braid (Douglas Hofstadter) — culture for Yoshua Bengio (wikidata.org)
- Causality: Models, Reasoning, and Inference by Judea Pearl (2000). Causality: Models, Reasoning, and Inference (Judea Pearl) — culture for Yoshua Bengio
- Finding Structure in Time by Jeffrey L. Elman (1990). Finding Structure in Time (Jeffrey L. Elman) — ghost for Yoshua Bengio
Peers and kindred spirits
- Yann LeCun. LeCun and Bengio co-founded a conference. (also via Learning Deep Architectures for AI, Backpropagation Applied to Handwritten Zip Code Recognition) “Yann is the co-director of the CIFAR program on Neural Computation and Adaptive Perception Program with Yoshua Bengio.” (Meta AI)
- Geoffrey Hinton. Hinton co-authored a letter with Bengio. (also via A Neural Probabilistic Language Model, Learning representations by back-propagating errors, ImageNet Classification with Deep Convolutional Neural Networks) “Sincerely, Yoshua Bengio Professor of Computer Science at Université de Montréal & Turing Award winner Geoffrey Hinton Emeritus Professor of Computer Science at University of Toronto & Turing Award winner” (Biocomm AI)
- Aaron Courville by Aaron Courville. The publication listing credits Bengio and Courville together on Deep Learning. (also via FiLM: Visual Reasoning with a General Conditioning Layer) “Ian Goodfellow, Yoshua Bengio and Aaron Courville: Deep Learning (Adaptive Computation and Machine Learning) , MIT Press, Cambridge (USA), 2016.” (Wikipedia)
- Ian Goodfellow by Ian Goodfellow. The book’s citation identifies Bengio and Goodfellow as coauthors. “title={Deep Learning}, author={Ian Goodfellow and Yoshua Bengio and Aaron Courville}, publisher={MIT Press}” (Deep Learning)
- Ilya Sutskever. Deep Learning (Yoshua Bengio) — geography for Ilya Sutskever (via Deep Learning, ImageNet Classification with Deep Convolutional Neural Networks, Sequence to Sequence Learning with Neural Networks) (openlibrary.org)
- International Conference on Learning Representations. Bengio co-founded the International Conference on Learning Representations with LeCun. “In 2013, he and Yoshua Bengio co-founded the International Conference on Learning Representations , which adopted a post-publication open review process he previously advocated on his website.” (Wikipedia)
- ImageNet Classification with Deep Convolutional Neural Networks by Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton (2012). ImageNet Classification with Deep Convolutional Neural Networks (Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton) — peer for Yoshua Bengio “Hinton received the 2018 Turing Award, together with Yoshua Bengio and Yann LeCun, for their work on deep learning.” (en.wikipedia.org)
- FiLM: Visual Reasoning with a General Conditioning Layer by Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin and Aaron Courville. The FiLM authors cite earlier work involving Bengio; the citation does not establish that Bengio collaborated on FiLM. “Dynamic Layer Norm for speech recognition [ \citeauthoryearKim, Song, and Bengio2017 ] , and Conditional Batch Norm for general visual question answering” (FiLM: Visual Reasoning with a General Conditioning Layer)
- Deep Learning in Neural Networks: An Overview by Jürgen Schmidhuber (2015). Deep Learning in Neural Networks: An Overview (Jürgen Schmidhuber) — peer for Yoshua Bengio “== Credit disputes == Schmidhuber has controversially argued that he and other researchers have been denied adequate recognition for their contribution to the field of deep learning, in favour of Geoffrey Hinton, Yoshua Bengio and Yann L…” (en.wikipedia.org)
- Martin Ford by Martin Ford. Bengio contributed a chapter to a book by Ford. “Bengio contributed one chapter to Architects of Intelligence: The Truth About AI from the People Building it , Packt Publishing, 2018, ISBN 978-1-78-913151-2 , by the American futurist Martin Ford .” (Wikipedia)
- Reinforcement Learning: An Introduction by Richard S. Sutton and Andrew Barto (1998). Reinforcement Learning: An Introduction (Richard S. Sutton and Andrew Barto) — geography for Yoshua Bengio (wikidata.org)
- Explaining and Harnessing Adversarial Examples by Ian Goodfellow, Jonathon Shlens, Christian Szegedy (2015). Explaining and Harnessing Adversarial Examples (Ian Goodfellow, Jonathon Shlens, Christian Szegedy) — collaborator for Yoshua Bengio (wikidata.org)
- The Neural Autoregressive Distribution Estimator by Hugo Larochelle and Iain Murray (2011). The Neural Autoregressive Distribution Estimator (Hugo Larochelle and Iain Murray) — collaborator for Yoshua Bengio (wikidata.org)
- Spoken Dialogues with Computers by Renato De Mori (1998). Spoken Dialogues with Computers (Renato De Mori) — geography for Yoshua Bengio (openlibrary.org)
- Fei-Fei Li. Gradient-based learning applied to document recognition (Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick Haffner) — peer for Fei-Fei Li (via Gradient-based learning applied to document recognition) “The same year, he received the grand prize of the VinFuture Prize alongside Yoshua Bengio, Jensen Huang, Geoffrey Hinton, and Fei-Fei Li for their groundbreaking contributions to neural networks and deep learning algorithms.” (en.wikipedia.org)
- Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning by Richard Sutton, Doina Precup, Satinder Singh (1999). Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning (Richard Sutton, Doina Precup, Satinder Singh) — geography for Yoshua Bengio
- FiLM: Visual Reasoning with a General Conditioning Layer by Ethan Perez, Florian Strub, Vincent Dumoulin, Aaron Courville (2018). FiLM: Visual Reasoning with a General Conditioning Layer (Ethan Perez, Florian Strub, Vincent Dumoulin, Aaron Courville) — collaborator for Yoshua Bengio
- Backpropagation Applied to Handwritten Zip Code Recognition by Yann LeCun, Bernhard Boser, John Denker et al. (1989). Backpropagation Applied to Handwritten Zip Code Recognition (Yann LeCun, Bernhard Boser, John Denker et al.) — peer for Yoshua Bengio
Who Yoshua Bengio influenced
- Attention Is All You Need by Ashish Vaswani, Noam Shazeer, Niki Parmar et al. (2017). Attention Is All You Need (Ashish Vaswani, Noam Shazeer, Niki Parmar et al.) — legacy for Yoshua Bengio (wikidata.org)
- Sequence to Sequence Learning with Neural Networks by Ilya Sutskever, Oriol Vinyals, Quoc Le (2014). Sequence to Sequence Learning with Neural Networks (Ilya Sutskever, Oriol Vinyals, Quoc Le) — legacy for Yoshua Bengio (wikidata.org)
- Playing Atari with Deep Reinforcement Learning by Volodymyr Mnih, Koray Kavukcuoglu, David Silver et al. (2013). Playing Atari with Deep Reinforcement Learning (Volodymyr Mnih, Koray Kavukcuoglu, David Silver et al.) — legacy for Yoshua Bengio (wikidata.org)