Artificial intelligence never ceases to surprise us, and not always for the better. This week, DeepSeek, a well-funded Chinese AI lab, launched its DeepSeek V3 model, designed for complex tasks like programming and writing.
The model has proven to be efficient and competitive in benchmark tests, but it has also shown a peculiar flaw: it is convinced that it is ChatGPT, the famous OpenAI chatbot.
The DeepSeek V3 “identity crisis”
Users on social networks and technology experts have shared tests where DeepSeek V3 not only presents itself as ChatGPT, but insists on being a version of GPT-4 released by OpenAI in 2023.
When asked for instructions related to DeepSeek, it responds with information about OpenAI APIs. He even tells the same jokes as GPT-4, replicating their punchlines.
This erroneous behavior has sparked both laughter and concern. As social media fills with memes about the model “identity crisis,” AI experts warn of issues related to model training and data integrity.
The origin of the problem
DeepSeek has not revealed many details about the data sources used to train V3. However, indications are that the model was exposed to texts generated by ChatGPT or GPT-4.
This could have happened accidentally, but it could also be a deliberate strategy to leverage knowledge from another model.
Mike Cook, an artificial intelligence researcher at King’s College London, compared this situation to making photocopies of a photocopy: “Every time we replicate something without the original source, we lose information and connection with reality.”
The result is a model that not only imitates the responses of the original, but also inherits its errors and biases.
Using data generated by ChatGPT could also violate OpenAI’s terms of service, which prohibit using its outputs to train models that compete with its products.
Needless to say, this issue, if proven true, raises legal questions about transparency and ethical practices in AI development.
A data contamination problem
The DeepSeek V3 case reflects a broader problem in the AI industry: data contamination. With the increasing use of generative models, the web is saturated with AI-generated content.
It is estimated that by 2026, 90% of online content could be created by automated systems. This situation makes it difficult to properly filter the data used to train new models.
When AI models absorb data generated by other models, there is a risk of creating a cycle of errors and biases.
Heidy Khlaaf, AI science director at the AI Now Institute, explained that developers might be tempted to resort to these practices to save costs, even if this compromises the quality of the model.
Industry reactions
The industry was quick to speak out. Sam Altman, CEO of OpenAI, posted on
On the other hand, DeepSeek is not the only model that has shown similar behaviors. Google Gemini, for example, has also come to be misidentified as other systems, such as Baidu’s Chinese chatbot Wenxinyiyan.
This suggests that the problem is not unique to DeepSeek, but rather a widespread consequence of training models with contaminated data.
Ethical and legal implications
The DeepSeek V3 incident highlights the urgency of establishing stricter regulations on the collection and use of data to train AI.
Without a clear legal framework, companies could face costly litigation and losses of public trust.
Furthermore, these types of incidents affect the reputation of the entire industry. If AI models cannot be trusted even to identify their own provenance, how will they gain trust in critical sectors such as healthcare or finance?
The future of AI after the incident
The controversy could accelerate the development of error mitigation technologies, such as Recovery Augmented Generation Verification (RAG-V).
These solutions seek to integrate verification steps that improve the accuracy and reliability of AI-generated responses.
However, technological solutions are not enough. It is essential that companies adopt ethical and transparent practices in model development. This includes:
- Review and clean training data.
- Respect the terms of use of other models.
- Implement rigorous quality controls before product launch.
Just the tip of an iceberg
The case of DeepSeek V3, which “self-perceives” itself as ChatGPT, is a reminder of the challenges and risks associated with competition in the field of AI.
We do not doubt that similar behaviors will come to light in the coming months, and there will even be talk of stagnation in the AI industry.
Although social networks may laugh at the situation, the legal, ethical and technological implications are serious. The AI industry must urgently address these issues to avoid further damage to its reputation and credibility.
This post is also available in: