New Model Reconstructs LLM Prompts with Near-Perfect Accuracy

TL;DR. Researchers developed a method to reconstruct large language model prompts from output text with high accuracy, posing a security risk for proprietary systems. - The "Previous-Token Prediction" inverse model trains on synthetic data to predict prior tokens without accessing model weights. - It can reconstruct exact prompts and semantic variants, even across different LLM architectures like Qwen and GPT-4o. - This technique threatens proprietary system prompts, potentially exposing trade secrets and moderation rules.

Sources

Back to QLANKR News