Advanced Statistical Methods for the Analysis of Large Data-Sets

Advanced Statistical Methods for the Analysis of Large Data-Sets

Many research studies in the social and economic fields regard the collection
and analysis of large amounts of data. These data sets vary in their nature and
complexity, they may be one-off or repeated, and they may be hierarchical, spatial,
or temporal. Examples include textual data, transaction-based data, medical data,
and financial time series.
Today most companies use IT to support all business automatic function; so
thousands of billions of digital interactions and transactions are created and carried
out by various networks daily. Some of these data are stored in databases; most
ends up in log files discarded on a regular basis, losing valuable information that is
potentially important, but often hard to analyze. The difficulties could be due to the
data size, for example thousands of variables and millions of units, but also to the
assumptions about the generation process of the data, the randomness of sampling
plan, the data quality, and so on. Such studies are subject to the problem of missing
data when enrolled subjects do not have data recorded for all variables of interest.
More specific problems may relate, for example, to the merging of administrative
data or the analysis of a large number of textual documents.
Standard statistical techniques are usually not well suited to manage this type
of data, and many authors have proposed extensions of classical techniques or
completely new methods. The huge size of these data sets and their complexity
require new strategies of analysis sometimes subsumed under the terms “data
mining” or “predictive analytics.” The inference uses frequentist, likelihood, or
Bayesian paradigms and may utilize shrinkage and other forms of regularization.
The statistical models are multivariate and are mainly evaluated by their capability
to predict future outcomes.



BAGIKAN
E-Book Lainnya