Analysis of 1,800 AI datasets: ~70% didn't state what license should be used or had been mislabeled with more permissive guidelines than their creators intended
I am very proud of this work. Data creation is already undervalued, but our current licensing means attribution is difficult even for practitioners who want to do the right thing. This is critical work to empower reliable data provenance and a call to improve data transparency.
📢Announcing the🌟Data Provenance Initiative🌟 🧭A rigorous public audit of 1800+ instruct/align datasets 🔍Explore/filter sources, creators & license conditions ⚠️We see a rising divide between commercially open v closed licensed data 🌐: https://dataprovenance.org/ 1/ [video]
wake up babe, the year's biggest data data set research project just dropped The Data Provenance Initiative analyzed 1,800+ popular fine-tuning text data sets and found a crisis of confusion. W/insights from @ShayneRedford @sarahookr https://www.washingtonpost.com/ ...
Context: A Crisis in Data Transparency ➡️Instruct/align finetuning often compiles 100s of datasets ➡️How can devs filter for datasets without legal/ethical risk, and understand the resulting data composition? 2/ [image]
Fantastic piece by @nitashatiku putting the spotlight on the release of the Data Provenance Initiative. We joined forces to audit and trace the data provenance of nearly 2,000 of the most widely used fine-tuning datasets. https://www.washingtonpost.com/ ...