Saw an AI training data audit at a local library last week
I stopped by the Brooklyn Public Library for a tech talk, and there was this volunteer going through old scanned books to flag biased language in the training set. She caught 12 examples in just one hour, stuff like gender-coded job titles from 1950s manuals. Honestly, it made me realize how much manual work goes into cleaning AI data, and I wondered how many libraries even do this kind of thing. Has anyone else seen grassroots audit efforts like this in their city?
Wait, is flagging old books even the best use of time when the real bias is probably baked into the search algorithms themselves? I mean, my local library has a whole section of WWII era cookbooks that literally say "for the little woman at home" and nobody's touched those in years. But hey, at least someone's doing something, right? I actually tried to volunteer for a similar thing at the Queens library last spring but they wanted me to sign a 40 page NDA first, which felt super weird for just reading old gardening manuals. Plus the volunteer coordinator kept calling it "data scrubbing" like we were cleaning toilets or something, which kind of killed the vibe for me.