Фундаментални публикацииA New Approach to Linear Filtering and Prediction Problems(1960) Kalman
Introduction to decision trees (1986)
Learning representations by back-propagating errors (1986)
Backpropagation applied to handwritten zip code recognition (1989)
Gradient-Based Learning Applied to Document Recognition (1998)
Long Short-Term Memory (1997)
A neural probabilistic language model (2003)
Anomaly Detection: A Survey (2009)
Causal Inference in Statistics: An Overview (2009)
ImageNet Classification with Deep Convolutional Neural Networks (2012)
Word2Vec: Distributed Representations of Words and Phrases and their Compositionality (2013)
Sequence to Sequence Learning with Neural Networks (2014)
Generative Adversarial Nets (2014)
Playing Atari with Deep Reinforcement Learning (2015)
Deep learning LeCun (2015)
Deep Learning for Time Series Classification: A Review
Attention Is All You Need (2017)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2018)
Data Movement is All You Need: A Case Study on Optimizing Transformers
FlashAttention: Fast and Memory- Efficient Exact Attention with IO-Awareness”
Онлайн курсовеhttps://www.coursera.org/specialization ... troductionhttps://www.coursera.org/specializations/deep-learninghttps://deeplearning.aihttps://www.fast.aihttps://www.youtube.com/live/r0Ogt-q956I?feature=sharedAndrej Karpathy Building GPT
https://www.youtube.com/playlist?list=P ... GvCAUhRvKZAndrej Karpathy MinGPT and NanoGPT
https://github.com/karpathy/nanoGPThttps://github.com/karpathy/minGPTKiril Aramenko - transformer architecture
https://community.superdatascience.com/c/llm-gptGrant Sanderson - visual math
https://www.youtube.com/playlist?list=P ... x_ZCJB-3pi Книгиhttps://www.statlearning.comhttps://fleuret.org/public/lbdl.pdf