BuzzFeed Tech Talks is BuzzFeed's speaker series where we gather with the rest of the NYC tech community and discuss topics that matter to all of us. The theme of this event is about building data products. Every company has data, but how do you actually use it? At BuzzFeed, data fuels the how and why behind most of content and products you see. Come hear about how the NYC community uses data to bring new features and ideas to life! SPEAKERS: Meghan Heintz is a Senior Data Scientist at BuzzFeed, primarily working on Tasty. That's right, you can blame her for those cheese stuffed top down cooking videos that thwart your Paleo diet plans. She previously worked at Zynga on dynamic difficulty tuning and predictive modeling. Before she was in data science, she worked as an environmental consultant on river and wetland restoration projects. She holds a B.S. in Environmental Resources Engineering from Humboldt State where she focused on fluid dynamic modeling. Lucy Wang is a Senior Data Scientist at BuzzFeed working on machine learning tools for optimizing audience reach and engagement. She holds an MS in Computer Science from Columbia University where she researched social networks and information diffusion. Her life peaked the day she got swarmed by baby pygmy goats. Greg Stoddard is a Research Director and data scientist at Crime Lab New York (https://urbanlabs.uchicago.edu/labs/crime-new-york) where he works on applying machine learning to problems in policing, criminal justice, and education. He one time tweeted something that got over 100 likes. He was proud of that. Dan Valente is a Machine Learning Engineering Manager at Spotify, where he's focused on leading and growing teams that deliver personalized musical experiences. Dan was formerly Director of Data Science at Chartbeat, where he built products that helped editorial teams understand reader behavior and deliver amazing content. When not pontificating about machine learning and products, Dan prefers to spend his time hanging out with his daughters and composing, listening to, talking about, or otherwise obsessing over music. PRESENTATIONS: Recipe2Vec: How word2vec Helped Us Discover Related Tasty Recipes Your user knows they want a healthyish but tasty pasta for dinner but aren't quite sure exactly which recipe to choose. How can you help narrow their search and show them closely related recipes to give them enough options without making their search exhausting? This talk will show you BuzzFeed/Tasty tech's solution to creating a consistent method for finding similar Tasty recipes using word2vec. Casting A Wide Neural Net For Content Discovery With hundreds of pieces of content produced daily, how do we make sure that every one of them reaches its full potential audience? We have a vast network of social media channels through which we distribute this content, so it's important that content makes it to all its relevant channels and doesn't slip through the cracks. Here I talk about a model we built to catch all content that fits the voice and audience of a particular channel -- a journey that involves deep learning, shallow learning, and human learning. Selectively Labeled Datasets: How To Know What You Don't Know Supervised learning assumes that an analyst has access to a representative dataset of features and labels for a given task. However in many applications, particularly in public policy, the dataset was generated by a series of unknown policies and hence may not be representative of the entire population. In this talk, I'll present a particular example of this phenomena, namely that the dataset may only have labels for a select subset of the population, and discuss strategies for addressing it. User Experience > Your Model When training algorithms on your petabytes of data, fighting for that 0.2% increase in accuracy, it is all too easy to forget one thing: you are building a product. And your product has users. Make no mistake, machine learning can help deliver some amazing user experiences, but it can also deliver horrible ones. Using Spotify's algorithmic playlists as examples, I'll lay out a few ideas to keep in mind when building ML-based products so that users are delighted, rather than repulsed, by the results of your algorithms --- including posing that most blasphemous question: should you use ML at all?
Type: Eventbrite