How social media outlets are being used to train AI

As AI grows stronger, new ways of training it have emerged. AI companies are turning to social media to train their models to be able to converse in a human manner and understand human-written questions. Reddit is considering legal action to prevent this, after agreeing an exclusive deal with OpenAI reportedly worth $60 million a year.

Reddit can be used as a way of training AI platforms that require to understand data about a variety of subjects. By using online forums, the AI models can learn to converse more naturally, by surveying the way people type and answer human questions. This is a crucial part of AI, with new technology making it increasingly difficult to differentiate between human and AI created content.

As Reddit prepares to bring a legal claim on AI companies, other social media platforms may follow. The reasons for this can be case dependent, with some companies looking to protect their user-created content as a form of data protection, and others looking to monetize their content and create a source of revenue.

There are two main infringements that could be used in future claims: copyright infringement, and breach of services. AI companies that use user created content could be infringing on copyright, making social media companies, and the user whose content is being used, able to produce a claim. AI companies may also be liable for breaching the Terms of Services of social media companies by taking and reproducing original content without permission. This could result in legal action being taken and an injunction being made. Taking someone’s own generated content could also produce problems in the Data Protection Act, providing AI companies with an additional legal risk.

Before agreeing to a licensing deal, The New York Times brought a case upon OpenAI and Microsoft. It was alleged that AI had used copyrighted material without permission, but most of the claims were dismissed by the courts. There has since been an agreement between the two parties for the access of archived, and current, articles.

Getty Images have also recently begun pursuing legal action against Stability AI over the use of their images being used to train AI, with the case being allowed to go to court.

One solution to this is for a licensing agreement between the AI companies and social media companies. Such agreements are being proposed by OpenAI, who have been actively engaging in talks with social media companies for exclusive rights to train their AI models. Reddit and Google have recently agreed on a deal, reported to be worth $60 million, to use Reddit’s user created content for training Google’s AI models.

Social media and AI companies reaching licensing agreements can see a new era of data protection policies, with companies having to decide whether to work with AI.

share this Article

Recent Articles

Written By: