Principles of Large Language Models (LLM)

P. O. Lysyi

Journal of Numerical and Applied Mathematics · 2025 · 인용 2

This paper explores the operational principles of large language models (LLMs), focusing in particular on the mechanism of next-token generation within the process of autoregressive modeling. It outlines the theoretical foundations of neural language models, the transformer architecture with its self-attention mechanism, and the roles of tokenization and embedding in forming the input representation of text. The study analyzes the main methods for selecting the next token (greedy decoding, top-k sampling, top-p sampling, temperature), their impact on the stochasticity of results, and the trade-off between coherence and creativity.

It also examines context length limitations, sources of training data, and challenges related to interpretability and the likelihood of «hallucinations». The article provides a comprehensive overview of the architectural and algorithmic foundations behind text generation in LLMs.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.