Probing Reveals Claude, GPT Training Timelines and Knowledge
TL;DR. New analysis uses 'Incompressible Knowledge Probes' and 'Data Mixture Inference' to estimate the pre-training timelines and data sources for models like GPT-5 and Claude Opus. - Researchers can approximate model parameters and dataset mixtures by carefully crafting requests to frontier LLMs. - Scoring models on date-related queries helps estimate the specific pre-training checkpoints and their recency. - Understanding training stages, from pre-training to post-training, provides insight into model capabilities and versions.
- Analysis uses probing techniques to deduce LLM pre-training timelines.
- Methodology approximates model parameters and dataset mixtures.
- Highlights the three main stages of large language model training.