Blog
193 posts · page 5 of 8
This post was published on Medium An important part of the data scientists and researchers’ life is to keep track of publications in their field. Depends on your field and needs…
This post was publish on Medium While there are many people who would like to become a data scientist and are looking for their first position, junior data science positions are…
Five Talks from spaCy-IRL Worth Watching - great summarisation of 5 talks from spaCy-IRL conference which took place in Berlin in the beginning of July. The summarisations are…
Checklist for debugging neural networks - well written trouble shooting for neural networks models which is not language or framework specific!…
- 4 insights from BDHW19·3 min
This week I attended BDHW19 - Big data in Health Care which was hosted by the Weizmann Institute of Science in collaboration with Nature Medicine. The conference had a great line…
- 5 > 11·1 min
But it takes 11 minutes to read the tutorial...
How to Grow Neat Software Architecture out of Jupyter Notebooks - jupyter notebooks is a very common tool used by data scientist. However, the gap between this code to production…
Deep density networks and uncertainty in recommender systems - Yoel Zeldes and Inbar Naor from Taboola engineering team published a series of posts (4 so far) about uncertainty in…
JQ cook book - I find myself using JQ quite often and sometimes to more complex things than just filtering fields. https://github.com/stedolan/jq/wiki/Cookbook Bonus point - list…
I have recently took "Bayesian machine learning in Python: A\B testing" course in Udemy . My notes from the course can be found here . It is mainly written for myself for easier…
https://mathwithbaddrawings.com/2017/11/29/epitaphs-in-the-graveyard-of-mathematics/ And a follow up - https://gorelik.net/2017/12/02/epitaphs-in-the-graveyard-of-mathematics/
Authors: David Tsurel , Dan Pelleg , Ido Guy , Dafna Shahaf Article can be found here Trivia facts can drive users engagement, But what are trivia fact? Is the fact “Barack Obama…
- Map Spark UDAF (Java)·1 min
I run Spark code on Java. I had data with the following schema - [code language="bash"] root |-- userId: string (nullable = true</span> |-- dt: string (nullable = true)</span> |--…
Code challenges are a common tool to evaluate candidate ability to develop software. Of course there are other indicators such as - blog posting, open source involvement, github…
- SO end of year surveys·3 min
Recently Stack Overflow published few posts comparing the usage of Stack Overflow between different segments \ scenarios: How Do Students Use Stack Overflow? What Programming…
- Davies-Bouldin Index·1 min
TL;DR - Yet another clustering evaluation metric Davies-Bouldin index was suggested by David L. Davies and Donald W. Bouldin in "A Cluster Separation Measure" ( IEEE Transactions…
Super Mario from Microsoft ( Daniel Molnar ) - Data Janitor 101 , one of the best reasoned talks I heard for a long time. Andrew Clegg , data scientist @ Etsy gave an historic…
[gallery ids="889,890" type="rectangular"] From Philipp Krenn's, Developer Advocate at Elastic, "Databases - The Choice is Yours" talk.
My summary and notes for "Detecting Data Errors: Where are we and what needs to be done?" by Ziawasch Abedjan, Xu Chu, Dong Deng, Raul Castro Fernandez, Ihab F. Ilyas, Mourad…
Joel test for Data Science - the inspiration and adjustment to data science done in Domino, both were interesting reads for me.…
- 5 Python NLP pacakges·2 min
NLP is a broad term which contains many types of question and challenges such as - language detection, Part-of-Speech tagging, relation extraction, named entity recognition, OCR,…
I got a diversity scholarship from Num Focus to attend the PyData Berlin event. Num Focus is an NGO which supports open source data science projects among them - Jupyter,…
Similar Wikipedia Pages - This post present Wikipedia similar pages chrome extension. Phrasing this in other words it is a recommendation system for Wikipedia pages. They are not…
Spoiler detector - Get your annotated data set for free! Nice way to get (though not perfect) an annotated data set and create some social good. Would be interesting to expand it…
My talk from Swiss Python Summit is online - https://www.youtube.com/watch?v=Q9AU_qETVd8