Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I think I agree with you overall — there's no way to perform this task perfectly for all time in such a way that it makes this problem disappear. But I want to push back against the specifics of your argument, namely about the cost and resources involved, as well as the lack of feedback mechanisms. I do that below; I wrote those paragraphs before really giving credit to how hard it would be to comprehensively score audience + topic pairs. That sounds like... exponentially growing complexity. So that approach sounds like a non-starter; maybe an alternative is to look at quality metrics on clusters of search queries as a whole, before diving into seeing if there are problems with individual audiences and those queries. Who knows.

There are many search queries to watch, but 'black girls' was persistently problematic for a long time. Noticing when a cluster of queries gets that status might be extremely resource intensive if that status was highly ephemeral; it's not. It's work that can be done asynchronously, and run daily at most. Is this simple? Nothing involving software is, imo. But it's probably much simpler than many other machine learning-driven pieces of Google's product.

Also, there are tons of feedback mechanisms available and actively used by Google. Every user interaction with search results is available to Google; a lot of these stand in for quality of result: did the person refine their query after seeing crappy results? Did they click a link and then press 'back' really quick? All of these factors already feature in Google's algorithm.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: