En tous cas, j'ai appris ça :
The Pangram AI detection tool operates differently from Turnitin’s AI detection tool. According to the tool’s technical documentation, Pangram was trained on approximately 28 million human-authored documents drawn from open-source datasets (Emi and Spero 2024). The aim is the same as Turnitin’s GenAI detection tool; to recognise human-written text and AI-generated text by LLMs. A feature of Pangram’s approach is the use of a “mirror prompt” technique. For each human-written text in their training data, they created a customised prompt to generate an AI version of that exact text, matching it in terms of topic, tone, style, and length. This approach is called “syntactic mirror” and creates AI-generated versions of human text that look almost the same in structure, style and flow. This makes it harder for the model to rely on obvious signs of GenAI writing and instead trains it to notice smaller, more subtle differences. An example concerns how sentences are built, which linking words are used, and how varied the writing is. By doing so, it pushes the model to look past metrics like perplexity or burstiness. Subsequently, Pangram employs another technique known as “hard negative mining”. In this process, the model searches for “difficult cases“ where it misclassified text or where the difference between human and GenAI writing was minimal. They used these cases to train the data further and identify the differences between GenAI and human writing. Pangram supports multilingual detection (Emi and Spero 2024).





