Showing posts with label Analyst. Show all posts
Showing posts with label Analyst. Show all posts

Monday, December 12, 2016

Applied Econometrics & Statistical Learning Algorithm: Random Trees and Random Forests

Author: Kateryna Volkovska

"... ensembles of decision trees - often known as Random Forests - have been the most successful general-purpose algorithm in the modern times." Howard & Bowles (2012)

In this blog post, I want to draw your attention to the very interesting and useful algorithm called Random Forest. In econometrics, Random Forests are used in GDP forecasting and poverty prediction. Also, this approach can be used to rank the importance of variables/classifier in regression and classification tasks (variable selection method). Variables that enter more trees/models (note that this share of trees often called the importance score) are stronger predictors, than those that enter fewer trees.

CART (Classification and regression trees)

Random Forest can contain not only hundreds but thousands and more of individual trees. That is why in order to understand the concept of random forest, we should firstly define what is the single random tree.
The main idea behind CART is quite simple: a set of observed predictors is used to recursively partition the data until the values of the respond variable become homogeneous with each sub- partition (Ikonen, 2016). CART splits one variable at a time. The best partitioning variable at each split is determined by minimizing the sum of squared errors in regression or alternatively finding the predictor that best splits the response variable into separate classes in classification (Sikory 2009). When the best available split is found, the procedure continues until the minimum node size is reached.

Random Forests (RF)

Random Forest (alternatively - Decision Forest) is the complex of decision trees (tree predictors) in which each individual tree is constructed from a unique subset of data with randomly selected observations and variables. My motivation for studying this algorithm more detailed was that it gives good results in classification and regression problems. For example. in the study by Caruana et al. (2008) indicated that Random Forests offer the most accurate and stable results.
Random Forests are known to be highly resistant to over-fitting and to effectively handle the noise (Payne, 2014).  Due to the majority voting, the algorithm also efficiently deals with the common problem in econometrics - heteroscedasticity.  Random Forests also expand upon the strengths of standard decision tree predictors in detecting non-linearity in the data and working with the high-dimensional data sets (Siroky, 2009; Caruana et al. (2008)).
Hare are the very intersting illustrations from the lecture slides of Cutler (2010) about how the Random Forest captures the data. Suppose we have the following data and underlying function:

That s how the single regression tree deals with data:
And if to take 10 regression trees:
And now the average of the 100 regression trees:
Let´s have a closer look at how the algorithm works.

How it works: intuition

The algorithm proposes to select randomly the subsets of the data and "grow" the decision tree on them independently.  At the last stage, it combines this decision trees and aggregates their predictors by majority voting, or simply, averaging.  Interesting explanation I have found in the Biau and Scornes (2010): the algorithm rules by "divide and conquer" principle - you should sample the small fractions of data, grow a randomized tree predictor on each small piece and then paste (aggregate) these predictors together. Thereby, the group of weak models combines to form a quite powerful model.
Now let´s look at my super-simple example in order to understand the intuition behind. Suppose Sally have her own criteria of an ideal boyfriend. The factors that influence her decision are age, whether he is smart or not, whether he is cool or not and whether he is handsome. Now we build three decision trees according to Sally´s preferences:
Then Sally meets a man (new object for the procedure) who has features:  20-30 years, handsome, cool but, unfortunately, not smart :(
The results from each tree are the following: decline, accept, decline. Therefore, using the majority rule (2 declines against 1 accept) - Sally will reject this guy :( So being smart is really important.

Poverty prediction  

Otok and Seftiana (2014) concluded that Random Forest performs nicely in defining poor households. At  the same time, Thoplan (2014) found that this algorithm predicts poverty quite accurate.
Now let´s study a more complex example. Sohnesen and Stender (2016) examined the problem of poverty and found that Random Forest often has higher accuracy and, in particular, good predicts in the rural/urban areas. In the example, they examined 6 countries: Albania, Ethiopia, Malawi, Rwanda, Tanzania and Uganda and compared the performance of Random Forest versus the other common tool for predicting poverty - Multiple Imputation (MI).
To do this they build 6 models of MI and RF with different variable selection methods and loss functions and compared the prediction accuracy:
  1. RF using Gini impurity function
  2. RF using Entropy loss function
  3. MI with Stepwise variable selection
  4. MI with LASSO variable selection
  5. MI with 25 variables based on the importance score from RF
  6. RF with 25 variables based on the importance score from RF.
Table 1. Mean squared error of poverty predictions for different variable selection methods and RF loss functions.

Source: Sohnesen and Stender (2016)  
Analysis of the forecasting performance of the linear regression based models an RF shows that both approaches do well at the national level (both patterns are fairly similar), but RF are in general more accurate. It has higher accuracy in 4 out of 6 countries and does better at the mean (Column 7 and 8 vs 3 and 4).
Sohnesen and Stender (2016) also conclude that RF is more robust and does not make large prediction errors at rural/urban levels unlike commonly applied linear regression models.

To summarize:

Advantages

  • The great pros of Random Forests are that the effects of  heteroscedasticity, outliers, and other data anomalies are reduced due to the large amount of individual tree learners and majority voting (Breiman, 2011).
  • The  algorithm provides unbiased estimates of model generalization error.
  • Excellently performs when the number of variables is much larger that the number of observations 
  • Very accurate approach and excellent classification algorithm. 
  • More robust predictor as it does not rely only on the one prediction model.
  • Particulary well-suited to the small sample size and large p-value problems.
  • Detect nonlinear relationships and good works with the high dimensional datasets (Siroky, 2009). 

Problems  

  • In big trees we can face the problem of over-fitting and, as the consequence, problems with generalizing the structure as the tree can be too much detailed.
  • In small trees we can be unable to capture the essential details in the data, some important specifications can be missed.  
  • Rather slow and mathematically complex algorithm.
  • So-called "black box": quite hard to get insights into decision rules.

Software & Languages

  • R (package randomForest: functions "randomForest" and "varlmpPlot") 
  • Python (package Anaconda)
Hope my post was interesting for you, have a great Christmas holidays and happy blogging! :)


Saturday, November 19, 2016

Boost your career: Interesting insights of Data Mining & Machine Learning for economists


 
 
 
With this post, I want to share some useful information and introduce you a couple of Machine Learning (ML) concepts as well as highlight differences and similarities between ML and econometrics. Also, I am aimed to show how Data Mining (DM) and ML can be used by economists.

Basics

So firstly let's clarify definitions:
  • ML uses data to predict some variable as a function of other variables (focuses on computing a good prediction of y given the new values of x). 
  • Econometrics uses statistical methods for prediction and inference of economic relationship.
In general, econometricians are thought to start with the theoretical model and then build a model that validates or invalidates the theory. Machine learners always start from data. 
Machine learning techniques (such as decision trees, support vector machines (SVM), neural networks and deep learning) allow for more effective ways to model complex economic relationships.
 
Table 1. The comparison of aims of Econometrics, DM and ML.

EconometricsMachine LearningData Mining
prediction
prediction
summarization
summarizationextract info from datafinding patterns
estimation

visualization
hypothesis testing

data manipulation
extract info from data


Source: Varian (2014)

So the main difference is that ML, for the most part, deals with pure prediction, while econometrics cares more on causal inference.

Predict & Classify

When econometrician faced up with prediction problem he or she usually employs the linear or logit regression. However, ML suggests more advanced nonlinear methods that are more useful for big data sets. Here are some of them:
  • Regression trees;
  • Random forest;
  • Least absolute shrinkage and selection operator (LASSO - regression analysis method);
  • Least-angle regression.
What economists always call “the out-of-sample prediction”, machine learners call “the case of overfitting”. The common difficulty for both is the classification problem. While econometrician in this case usually uses logit or probit, ML suggests using decision trees in order to classify the observation which will lead to good out-of-sample predictions (in literature you can find the abbreviation “CART” - classification and regression trees). The feature of decision trees is that they capture non-linearity in data, while logistic regression not. Hence, ML tool does better. Another cool thing of ML is that it prefers averaging over many small models which give better out-of-sample prediction than choosing a single model.

Data Structures  and Dimensionality Reduction

So what are the differences between data structures that are most commonly used by ML and econometrics? Firstly, econometricians deal usually with time-series and panel data, while machine learners prefer cross-sectional data with independent identically distributed observations. However, for time series ML offers a method called Bayesian structural time series (BSTS) aimed to work better for variable selection problems in time series application. 
I think all of you have heard about Principal component analysis (PSA) for dimensionality reduction of data. In fact, it is ML method, but it is widely used by econometricians and mathematicians. I used it also in my Bachelor thesis while making the analysis which particular factors influence most on the costs of insurance companies in the USA.

Regression Everywhere

Another common tool for machine learning specialists and econometricians is regression analysis. Its primary goal is to understand as far as possible with the available data, how the conditional distribution of the response y varies across subpopulation determined by the possible values of the predictors or predictor (Cook and Weisberg (1999)). In my post I want to catch your attention on the following economic example which provides methods for variable selection in the context of the growth regressions (Varian 2014).
In the example, he uses the dataset from Sala-i-Martin (1997) of 72 countries and 42 variables in order to determine the most important variables for economic growth. Sala-i-Martin (1997) computed all possible subsets of regressors and used the results to construct the measure called CDF(0). In the table below you can see the variables that have the highest CDF(0) and therefore the most useful in explaining economic growth according to Sala-i-Martin (1997). Ley and Steel (2009) for this problem used Bayesian model averaging, LASSO and spike-and-slab regressions (which is also a Bayesian technique) . In the following table LASSO column shows the ordinal importance of the variables or a dash meaning that it was not included in the chosen model. Other columns show the posterior probability of inclusion in the model.
 
Table 2. Comparing Variable Selection Algorithms: Which Variables Appeared as Important Predictors of Economic Growth?
 

Source: Ley and Steel (2009), data from Sala-i-Martin (1997).
These methods are efficient and useful for economic research in case you faced up with the problem of determining variables that are most important for the particular model.

Must-have Software

So what about software and programming languages? For ML it is definitely R and Python (check packages “scikit learn” and “statsmodels”). And for econometrics R, Stata and Eviews are the best. The last two are the statistical software and they are not for free, thus, on my opinion, R is the most suitable for both purposes. For those, who are interested, I highly encourage to read a book “An Introduction to Statistical Learning” by Gareth James (https://www.amazon.com/Introduction-Statistical-Learning-Applications-Statistics/dp/1461471370).

Conclusion and Further Inspiration

We can see that econometrics and ML are very closely related. However, I consider econometrics as a subpart of ML. Other important applications of ML include:
  • Computer vision; 
  • Speech recognition (e.g. Siri and Hello Google that you all know); 
  • Artificial intelligence (check game “Just dance”:D)
To summarize, what the econometric community can learn from the ML community:
  • Tests to avoid overfitting; 
  • Nonlinear estimations; 
  • Model averaging; 
  • Tools for manipulating big data (SQL, NoSQL databases); 
  • Computational Bayesian methods.
I am convinced that ML tools should be more widely known by young economists and researchers. Hope you have found from this post some interesting ideas what you should learn to grow more in your future career.
And I want to share with you this beautiful mind map which provides the great overview of all machine learning techniques: http://machinelearningmastery.com/a-tour-of-machine-learning-algorithms/

Take blanket and cup of tea and start watching Data Mining video lectures of Jeef Leek (https://www.youtube.com/user/jtleek2007)

I would be very happy for your feedback and comments. Let's share ideas!

Happy blogging!:)