KyndaIlya Sutskever
InfluencesPeers & partnersSuccessors
Hover for the evidence (or the center for a bio) · click to travel · drag a bubble to tug it · scroll to zoom
Drawing the map…

Ilya Sutskever: influences, peers and legacy

Every connection with its receipt

Ilya Sutskever's intellectual lineage runs from Toronto's connectionist revival through information theory and into the scaling era he helped define. This mix traces the backpropagation and Boltzmann-machine inheritance of Geoffrey Hinton's lab, the algorithmic-information theorists who underwrite his "compression is intelligence" framing, and the papers that carried his sequence models into the transformer age.

The Kynda mix for Ilya Sutskever

  • Key Influence · A Fast Learning Algorithm for Deep Belief Nets by Geoffrey Hinton (2006). Hinton was Sutskever's doctoral advisor at the University of Toronto, and this paper is the one that restarted deep learning as a research program: greedy layer-wise pretraining of stacked restricted Boltzmann machines made deep nets trainable again after the 1990s winter. Sutskever arrived in Hinton's group in exactly this period, co-authored with him through AlexNet, and then co-founded Google Brain work and OpenAI from that foundation.
  • Influencia Obscura · A Formal Theory of Inductive Inference by Ray Solomonoff (1964). Solomonoff's theory of universal induction — predict by weighting all programs consistent with the data, shorter ones more heavily — is the formal skeleton behind Sutskever's much-discussed argument that next-token prediction and lossless compression are routes to general intelligence. In talks on unsupervised learning he has leaned explicitly on Kolmogorov-complexity and compression arguments of exactly this lineage.
  • Local Roots · Bayesian Learning for Neural Networks by Radford Neal (1996). Neal's Toronto thesis, supervised by Hinton, brought Markov chain Monte Carlo and Hamiltonian dynamics to neural network inference and remains the canonical Bayesian neural net text. It maps the local intellectual terrain Sutskever entered: a single department where Boltzmann machines, MCMC and variational methods were everyday tools, and where the Hessian-free optimization work he did with James Martens also grew.
  • Beyond the Medium · A Mathematical Theory of Communication by Claude Shannon (1948). Shannon founded information theory partly by modelling English as a Markov source and measuring the entropy of predicted next characters — literally a language model, with human subjects guessing letters. Cross-entropy loss, perplexity and the compression arguments Sutskever uses to justify next-token prediction all descend from this paper, making a 1940s telecom engineering document the hidden basis of GPT-era objectives.
  • Peer · Generating Sequences With Recurrent Neural Networks by Alex Graves (2013). Graves, another Hinton-adjacent researcher who moved to DeepMind, pushed LSTM sequence generation for text and handwriting in the same window Sutskever was building character-level RNNs and then sequence-to-sequence translation. The two represent parallel attacks on the same problem — making recurrent nets actually generate structured sequences — before attention and transformers absorbed both lines.
  • From the Canon · Sequence to Sequence Learning with Neural Networks by Ilya Sutskever (2014). Written with Oriol Vinyals and Quoc Le at Google, this is the paper that defined the encoder-decoder paradigm: read a source sentence into a vector with one LSTM, decode a translation with another, reverse the input word order for better gradients. It moved deep learning from fixed-size classification to arbitrary sequence mapping, and is the most direct ancestor of every later text-generation system.
  • Key Collaborator · Learning Multiple Layers of Features from Tiny Images by Alex Krizhevsky (2009). Krizhevsky's technical report introduced CIFAR-10 and CIFAR-100 and the GPU-minded engineering practice that made AlexNet possible three years later. He was the hands-on implementer of the network Sutskever and Hinton co-authored; this document is the groundwork, the dataset curation and layer-wise feature study that preceded their shared ImageNet result and the DNNresearch acquisition by Google.
  • Legacy · Attention Is All You Need by Ashish Vaswani (2017). The transformer paper is explicitly framed against the recurrent encoder-decoder models Sutskever, Vinyals and Le introduced, citing seq2seq as the dominant approach it replaces by dropping recurrence for attention alone. It inherits the problem statement, the translation benchmarks and the encoder-decoder shape wholesale — and the GPT series Sutskever then led at OpenAI took it straight back up.

What influenced Ilya Sutskever

  • Geoffrey Hinton. Sutskever is identified among Hinton's former students or postdoctoral researchers. (also via Attention Is All You Need, A Fast Learning Algorithm for Deep Belief Nets) “His doctoral advisor was Geoffrey Hinton.” (Wikipedia)
  • A Fast Learning Algorithm for Deep Belief Nets by Geoffrey Hinton (2006). A Fast Learning Algorithm for Deep Belief Nets (Geoffrey Hinton) — titan for Ilya Sutskever “With Alex Krizhevsky and Geoffrey Hinton, he co-created AlexNet, a convolutional neural network.” (en.wikipedia.org)
  • Computing Machinery and Intelligence by Alan Turing (1950). Computing Machinery and Intelligence (Alan Turing) — culture for Ilya Sutskever (wikidata.org)
  • Universal Artificial Intelligence by Marcus Hutter (2005). Universal Artificial Intelligence (Marcus Hutter) — ghost for Ilya Sutskever (openlibrary.org)
  • Parallel Distributed Processing by David Rumelhart (1986). Parallel Distributed Processing (David Rumelhart) — titan for Ilya Sutskever (openlibrary.org)
  • Fei-Fei Li. ImageNet Classification with Deep Convolutional Neural Networks (AlexNet) (Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton) — legacy for Fei-Fei Li (via ImageNet Classification with Deep Convolutional Neural Networks (AlexNet)) “When the New York Times reporter Cade Metz asked Hinton to explain in simpler terms how the Boltzmann machine could "pretrain" backpropagation networks, Hinton quipped that Richard Feynman reportedly said: "Listen, buddy, if I could expl…” (en.wikipedia.org)
  • Gradient-Based Learning Applied to Document Recognition by Yann LeCun (1998). Gradient-Based Learning Applied to Document Recognition (Yann LeCun) — titan for Ilya Sutskever
  • A Mathematical Theory of Communication by Claude Shannon (1948). A Mathematical Theory of Communication (Claude Shannon) — culture for Ilya Sutskever
  • A Formal Theory of Inductive Inference by Ray Solomonoff (1964). A Formal Theory of Inductive Inference (Ray Solomonoff) — ghost for Ilya Sutskever
  • Receptive Fields of Single Neurones in the Cat's Striate Cortex by David Hubel (1959). Receptive Fields of Single Neurones in the Cat's Striate Cortex (David Hubel) — culture for Ilya Sutskever
  • Three Approaches to the Quantitative Definition of Information by Andrey Kolmogorov (1965). Three Approaches to the Quantitative Definition of Information (Andrey Kolmogorov) — ghost for Ilya Sutskever

Peers and kindred spirits

  • OpenAI. Sutskever co-founded OpenAI. “OpenAI’s research director is Ilya Sutskever, one of the world experts in machine learning. Our CTO is Greg Brockman, formerly the CTO of Stripe.” (Future of Life Institute)
  • Alex Krizhevsky. Sutskever collaborated with Krizhevsky on AlexNet. (also via Learning Multiple Layers of Features from Tiny Images) “ImageNet Classification with Deep Convolutional Neural Networks Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton” (NeurIPS Proceedings)
  • Oriol Vinyals. Sutskever worked with Vinyals on a learning algorithm. (also via Show and Tell: A Neural Image Caption Generator) “Title: Sequence to Sequence Learning with Neural Networks Authors: Ilya Sutskever , Oriol Vinyals , Quoc V. Le” (arXiv)
  • Andrew Ng. Sutskever held a postdoctoral position with Ng. “And before that, I was a postdoc in Stanford with Andrew Ng 's group.” (Ilya Sutskever's home page)
  • Daniel Gross. Sutskever founded the company with Gross. “As you know, Daniel Gross’s time with us has been winding down, and as of June 29 he is officially no longer a part of SSI. We are grateful for his early contributions to the company and wish him well in his next endeavor.” (Safe Superintelligence Inc.)
  • Quoc Viet Le. Sutskever worked with Le on a learning algorithm. “Title: Sequence to Sequence Learning with Neural Networks Authors: Ilya Sutskever , Oriol Vinyals , Quoc V. Le” (arXiv)
  • Google Brain. Google Brain hired Sutskever as a research scientist. “I spent three wonderful years as a Research Scientist at the Google Brain Team.” (Ilya Sutskever's home page)
  • DNNResearch. Sutskever joined DNNResearch. “Before that, I was a co-founder of DNNresearch .” (Ilya Sutskever's home page)
  • Daniel Levy. Sutskever founded the company with Levy. “I am now formally CEO of SSI, and Daniel Levy is President. The technical team continues to report to me.” (Safe Superintelligence Inc.)
  • Yoshua Bengio. ImageNet Classification with Deep Convolutional Neural Networks (Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton) — peer for Yoshua Bengio (via ImageNet Classification with Deep Convolutional Neural Networks, Sequence to Sequence Learning with Neural Networks, Deep Learning) “Hinton received the 2018 Turing Award, together with Yoshua Bengio and Yann LeCun, for their work on deep learning.” (en.wikipedia.org)
  • DNNresearch Inc.. Sutskever co-founded DNNresearch with Hinton. “He co-founded DNNresearch Inc. in 2012 with his two graduate students, Alex Krizhevsky and Ilya Sutskever” (Wikipedia)
  • Safe Superintelligence Inc.. Sutskever founded Safe Superintelligence Inc. “In June 2024, Sutskever announced Safe Superintelligence Inc. (SSI), a new company he founded with Daniel Gross and Daniel Levy with offices in Palo Alto and Tel Aviv .” (Wikipedia)
  • Show and Tell: A Neural Image Caption Generator by Oriol Vinyals (2015). Show and Tell: A Neural Image Caption Generator (Oriol Vinyals) — collaborator for Ilya Sutskever “At Google Brain, Sutskever worked with Oriol Vinyals and Quoc Viet Le to create the sequence-to-sequence learning algorithm, and worked on TensorFlow.” (en.wikipedia.org)
  • Deep Learning by Yoshua Bengio (2016). Deep Learning (Yoshua Bengio) — geography for Ilya Sutskever (openlibrary.org)
  • Bayesian Learning for Neural Networks by Radford Neal (1996). Bayesian Learning for Neural Networks (Radford Neal) — geography for Ilya Sutskever (wikidata.org)
  • Unsupervised Representation Learning with Deep Convolutional GANs by Alec Radford (2015). Unsupervised Representation Learning with Deep Convolutional GANs (Alec Radford) — collaborator for Ilya Sutskever “This development made OpenAI chief scientist Ilya Sutskever consider that a future model, using more diverse language data, could map far more structures of meaning, eventually becoming a "learned core module" for superintelligence.” (en.wikipedia.org)
  • Learning Multiple Layers of Features from Tiny Images by Alex Krizhevsky (2009). Learning Multiple Layers of Features from Tiny Images (Alex Krizhevsky) — collaborator for Ilya Sutskever “With Alex Krizhevsky and Geoffrey Hinton, he co-created AlexNet, a convolutional neural network.” (en.wikipedia.org)
  • Generating Sequences With Recurrent Neural Networks by Alex Graves (2013). Generating Sequences With Recurrent Neural Networks (Alex Graves) — peer for Ilya Sutskever
  • Deep Boltzmann Machines by Ruslan Salakhutdinov (2009). Deep Boltzmann Machines (Ruslan Salakhutdinov) — geography for Ilya Sutskever
  • Neural Machine Translation by Jointly Learning to Align and Translate by Dzmitry Bahdanau (2014). Neural Machine Translation by Jointly Learning to Align and Translate (Dzmitry Bahdanau) — peer for Ilya Sutskever
  • Human-level control through deep reinforcement learning by Volodymyr Mnih (2015). Human-Level Control Through Deep Reinforcement Learning (Volodymyr Mnih) — peer for Ilya Sutskever

Who Ilya Sutskever influenced

  • Attention Is All You Need by Ashish Vaswani (2017). Attention Is All You Need (Ashish Vaswani) — legacy for Ilya Sutskever “When the transformer came out, literally as soon as the paper came out, literally the next day, it was clear to me, to us, that transformers addressed the limitations of recurrent neural networks, of learning long-term dependencies.” (HackerNoon)
  • Scaling Laws for Neural Language Models by Jared Kaplan (2020). Scaling Laws for Neural Language Models (Jared Kaplan) — legacy for Ilya Sutskever