Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Wednesday, November 15, 2023

Data Science - Books - Bibliography - Introduction

 

https://towardsdatascience.com/learn-on-towards-data-science-52245bc91451


Practitioner’s Guide to Data Science

By Hui Lin, Ming Li

1st Edition

First Published 2023

eBook Published 24 May 2023




Based on industry experience, this book outlines real-world scenarios and discusses pitfalls that data science practitioners should avoid. It also covers the big data cloud platform and the art of data science, such as soft skills. The authors use R as the primary tool and provide code for both R and Python. 

This book is for readers who want to explore possible career paths and eventually become data scientists. This book comprehensively introduces various data science fields, soft and programming skills in data science projects, and potential career paths. Traditional data-related practitioners such as statisticians, business analysts, and data analysts will find this book helpful in expanding their skills for future data science careers. Undergraduate and graduate students from analytics-related areas will find this book beneficial to learn real-world data science applications. Non-mathematical readers will appreciate the reproducibility of the companion R and python codes.


Key Features:

• It is hands-on. We provide the data and repeatable R and Python code in notebooks. Readers can repeat the analysis in the book using the data and code provided. We also suggest that readers modify the notebook to perform analyses with their data and problems, if possible. The best way to learn data science is to do it!



TABLE OF CONTENTS

Chapter 1|28 pages

Introduction

 

Chapter 2|18 pages

Soft Skills for Data Scientists

 

Chapter 3|8 pages

Introduction to the Data

 

Chapter 4|22 pages

Big Data Cloud Platform

 

Chapter 5|26 pages

Data Pre-processing

 

Chapter 6|22 pages

Data Wrangling

 

Chapter 7|26 pages

Model Tuning Strategy

 

Chapter 8|16 pages

Measuring Performance

 

Chapter 9|20 pages

Regression Models

 

Chapter 10|30 pages

Regularization Methods

 

Chapter 11|42 pages

Tree-Based Methods

 

Chapter 12|78 pages

Deep Learning

 


https://linhui.org/hui's_files/datascientist1#(20)

https://scholar.google.com/citations?user=PAArLQIAAAAJ&hl=en&oi=sra

https://scholar.google.com/citations?user=PAArLQIAAAAJ&hl=en

https://linhui.org/

https://github.com/happyrabbit

https://scientistcafe.com/

A Tour of Data Science: Learn R and Python in Parallel

Nailong Zhang

CRC Press, 11-Nov-2020 - Computers - 216 pages (C) 2021.

A Tour of Data Science: Learn R and Python in Parallel covers the fundamentals of data science, including programming, statistics, optimization, and machine learning in a single short book. It does not cover everything, but rather, teaches the key concepts and topics in Data Science. It also covers two of the most popular programming languages used in Data Science, R and Python, in one source.

Key features:

Allows you to learn R and Python in parallel

Cover statistics, programming, optimization and predictive modelling, and the popular data manipulation tools – data table and pandas

Provides a concise and accessible presentation

Includes machine learning algorithms implemented from scratch, linear regression, lasso, ridge, logistic regression, gradient boosting trees, etc.

Appealing to data scientists, statisticians, quantitative analysts, and others who want to learn programming with R and Python from a data science perspective.

A Hands-On Introduction to Data Science

Chirag Shah

Cambridge University Press, 02-Apr-2020 - Business & Economics - 424 pages


This book introduces the field of data science in a practical and accessible manner.

The foundational ideas and techniques of data science are provided  allowing students to easily develop a firm understanding of the subject. The material that will have continual relevance even after tools and technologies change. 

Using popular data science tools such as Python and R, the book offers many examples of real-life applications, with practice ranging from small to big data. A suite of online material for both instructors and students provides a strong supplement to the book, including datasets, chapter slides, solutions, sample exams and curriculum suggestions. This entry-level textbook is ideally suited to readers from a range of disciplines wishing to build a practical, working knowledge of data science.

https://books.google.co.in/books?id=rljPDwAAQBAJ

Data Science Job: How to become a Data Scientist

Przemek Chojecki, 31-Jan-2020 - Computers - 100 pages

Data Scientist is one of the hottest job on the market right now. Demand for data science is huge and will only grow, and it seems like it will grow much faster than the actual number of data scientists. So if you want to make a career change and become a data scientist, now is the time.

This book will guide you through the process. From my experience of working with multiple companies as a project manager, a data science consultant or a CTO, I was able to see the process of hiring data scientists and building data science teams. I know what’s important to land your first job as a data scientist, what skills you should acquire, what you should show during a job interview.

https://books.google.co.in/books?id=h0PZDwAAQBAJ


Foundations of Data Science

Avrim Blum, John Hopcroft, Ravindran Kannan

Cambridge University Press, 23-Jan-2020 - Computers - 432 pages

This book provides an introduction to the mathematical and algorithmic foundations of data science, including machine learning, high-dimensional geometry, and analysis of large networks. 

Topics include the counterintuitive nature of data in high dimensions, important linear algebraic techniques such as singular value decomposition, the theory of random walks and Markov chains, the fundamentals of and important algorithms for machine learning, algorithms and analysis for clustering, probabilistic models for large networks, representation learning including topic modelling and non-negative matrix factorization, wavelets and compressed sensing. 

Important probabilistic techniques are developed including the law of large numbers, tail inequalities, analysis of random projections, generalization guarantees in machine learning, and moment methods for analysis of phase transitions in large random graphs. Additionally, important structural and complexity measures are discussed such as matrix norms and VC-dimension. This book is suitable for both undergraduate and graduate courses in the design and analysis of algorithms for data.

https://books.google.co.in/books?id=koHCDwAAQBAJ


Data Science and Intelligent Applications: Proceedings of ICDSIA 2020

Ketan Kotecha, Vincenzo Piuri, Hetalkumar N. Shah, Rajan Patel

Springer Nature, 17-Jun-2020 - Technology & Engineering - 576 pages

This book includes selected papers from the International Conference on Data Science and Intelligent Applications (ICDSIA 2020), hosted by Gandhinagar Institute of Technology (GIT), Gujarat, India, on January 24–25, 2020. The proceedings present original and high-quality contributions on theory and practice concerning emerging technologies in the areas of data science and intelligent applications. The conference provides a forum for researchers from academia and industry to present and share their ideas, views and results, while also helping them approach the challenges of technological advancements from different viewpoints.


The contributions cover a broad range of topics, including: collective intelligence, intelligent systems, IoT, fuzzy systems, Bayesian networks, ant colony optimization, data privacy and security, data mining, data warehousing, big data analytics, cloud computing, natural language processing, swarm intelligence, speech processing, machine learning and deep learning, and intelligent applications and systems. Helping strengthen the links between academia and industry, the book offers a valuable resource for instructors, students, industry practitioners, engineers, managers, researchers, and scientists alike.

p.217 Human activity recognition

https://books.google.co.in/books?id=eSbsDwAAQBAJ


© 2020

Data Science and Productivity Analytics

Editors: Charles, Vincent, Aparicio, Juan, Zhu, Joe (Eds.)


Table of contents (15 chapters)

Data Envelopment Analysis and Big Data: Revisit with a Faster Method Pages 1-34

Khezrimotlagh, Dariush (et al.)

Data Envelopment Analysis (DEA): Algorithms, Computations, and Geometry Pages 35-56

Dulá, José H.

An Introduction to Data Science and Its Applications  Pages 57-81

Rabasa, Alex (et al.)

Identification of Congestion in DEA Pages 83-119

Mehdiloo, Mahmood (et al.)

Data Envelopment Analysis and Non-parametric Analysis Pages 121-160

Villa, Gabriel (et al.)

The Measurement of Firms’ Efficiency Using Parametric Techniques Pages 161-199

Orea, Luis

Fair Target Setting for Intermediate Products in Two-Stage Systems with Data Envelopment Analysis

Pages 201-226

An, Qingxian (et al.)

Fixed Cost and Resource Allocation Considering Technology Heterogeneity in Two-Stage Network Production Systems Pages 227-249

Ding, Tao (et al.)

Efficiency Assessment of Schools Operating in Heterogeneous Contexts: A Robust Nonparametric Analysis Using PISA 2015 Pages 251-277

Cordero, Jose Manuel (et al.)

A DEA Analysis in Latin American Ports: Measuring the Performance of Guayaquil Contecon Port 

Pages 279-309

Morales-Núñez, Emilio J. (et al.)

Effects of Locus of Control on Bank’s Policy—A Case Study of a Chinese State-Owned Bank 

Pages 311-335

Xu, Cong (et al.)

A Data Scientific Approach to Measure Hospital Productivity Pages 337-358

Daneshvar Rouyendegh (B. Erdebilli), Babak (et al.)

Environmental Application of Carbon Abatement Allocation by Data Envelopment Analysis Pages 359-389

Yu, Anyu (et al.)

Pension Funds and Mutual Funds Performance Measurement with a New DEA (MV-DEA) Model Allowing for Missing Variables Pages 391-413

Badrizadeh, Maryam (et al.)

Sharpe Portfolio Using a Cross-Efficiency Evaluation Pages 415-439

Landete, Mercedes (et al.)

https://www.springer.com/gp/book/9783030433833



Special Issue on Data Science for Better Productivity

Data science for better productivity

Vincent Charles,Juan Aparicio &Joe Zhu 

Journal of the Operational Research Society 

Volume 72, 2021 - Issue 5: Special Issue Data Science for Better Productivity



Afsharian, M. (2019). A frontier-based facility location problem with a centralised view of measuring the performance of the network. Journal of the Operational Research Society, 72(5), 1058–1074. https://doi.org/10.1080/01605682.2019.1639476   

Bougnol, M.-L., & Dulà, J. (2020). Improving productivity using government data: The case of US Centers for Medicare & Medicaid's ‘Nursing Home Compare. Journal of the Operational Research Society, 72(5), 1075–1086. https://doi.org/10.1080/01605682.2020.1724056   

Del Vecchio, M., Kharlamov, A., Parry, G., & Pogrebna, G. (2020). Improving productivity in Hollywood with data science: Using emotional arcs of movies to drive product and service innovation in entertainment industries. Journal of the Operational Research Society, 72(5), 1110–1137. https://doi.org/10.1080/01605682.2019.1705194   

Grimaldi, D., Fernandez, V., & Carrasco, C. (2019). Exploring data conditions to improve business performance. Journal of the Operational Research Society, 72(5), 1087–1098. https://doi.org/10.1080/01605682.2019.1590136   

Ihrig, S., Ishizaka, A., Brech, C., & Fliedner, T. (2019). A new hybrid method for the fair assignment of productivity targets to indirect corporate processes. Journal of the Operational Research Society, 72(5), 989–1001. https://doi.org/10.1080/01605682.2019.1639477   

Jiang, R., Yang, Y., Chen, Y., & Liang, L. (2019). Corporate diversification, firm productivity and resource allocation decisions: The data envelopment analysis approach. Journal of the Operational Research Society, 72(5), 1002–1014. https://doi.org/10.1080/01605682.2019.1568841   

Li, Y., & Chen, W. (2019). Entropy method of constructing a combined model for improving loan default prediction: A case study in China. Journal of the Operational Research Society, 72(5), 1099–1109. https://doi.org/10.1080/01605682.2019.1702905   

Lin, S.-W., Lu, W.-M., & Lin, F. (2020). Entrusting decisions to the public service pension fund: An integrated predictive model with additive network DEA approach. Journal of the Operational Research Society, 72(5), 1015–1032. https://doi.org/10.1080/01605682.2020.1718011   

Routh, P., Roy, A., & Meyer, J. (2020). Estimating customer churn under competing risks. Journal of the Operational Research Society, 72(5), 1138–1155. https://doi.org/10.1080/01605682.2020.1776166   

Shi, Y., Zhu, J., & Charles, V. (2020). Data science and productivity: A bibliometric review of data science applications and approaches in productivity evaluations. Journal of the Operational Research Society, 72(5), 975–988. https://doi.org/10.1080/01605682.2020.1860661   

Summerfield, N. S., Deokar, A. V., Xu, M., & Zhu, W. (2020). Should drivers cooperate? Performance evaluation of cooperative navigation on simulated road networks using network DEA. Journal of the Operational Research Society, 72(5), 1042–1057. https://doi.org/10.1080/01605682.2019.1700766   

Zhu, J. (2020). DEA under big data: Data enabled analytics and network data envelopment analysis. Annals of Operations Research, 1–23. In press. https://doi.org/10.1007/s10479-020-03668-8 

Zhu, W., Liu, B., Lu, Z., & Yu, Y. (2020). A DEALG methodology for prediction of effective customers of internet financial loan products. Journal of the Operational Research Society, 72(5), 1033–1041. https://doi.org/10.1080/01605682.2019.1700188 [Taylor & Francis On 

https://www.tandfonline.com/doi/full/10.1080/01605682.2021.1892466



Ud. 16.11,2023, 3.45 am Austin, Texas

Pub. 16.7.2021














What is Data Science? - An Introduction to Data Science - New Developments


What is Data Science? - An Introduction to Data Science


Data driven or data analysis driven decision making is age old. But new data processing technology allows people to process data in ways that was not done before. Hence data will drive business decisions much more intensively in the next decade.


IT departments are not content anymore with just providing technology for processing data. The discipline and the profession of  IT is getting  involved in finding and understanding the relevance of new data sources, big and small.

The practice of business intelligence is  expanding to create to develop capabilities for analyzing and visualizing structured and unstructured data for their relevance for business decision making, and then building applications that can be run on a periodic basis which can be as small as even seconds to take crime or fraud prevention activities.

Data science is the name of this emerging discipline.

Data Science Tutorial 1 - Video

__________________________

__________________________
edureka!

More videos are available on YouTube on Data Science




Concise Visual Summary of Deep Learning Architectures
Basically neural network architectures
http://www.datasciencecentral.com/profiles/blogs/concise-visual-summary-of-deep-learning-architectures


http://www.datasciencecentral.com has number of articles on data science.


Data Science - New Developments

2023



50 Years of Data Science
David Donoho
Journal of Computational and Graphical Statistics
Volume 26, 2017 - Issue 4
Pages 745-766  Published online: 19 Dec 2017
https://www.tandfonline.com/doi/full/10.1080/10618600.2017.1384734

2020
The 2020 Data Science Dictionary—Key Terms You Need to Know
https://www.datasciencecentral.com/profiles/blogs/top-data-science-skills-for-2020-1

Trends in Artificial Intelligence and Data Science for 2020
https://www.datasciencecentral.com/profiles/blogs/trends-in-artificial-intelligence-and-data-science-for-2020-by

Top 5 Data Science Trends for 2020
https://www.datasciencecentral.com/profiles/blogs/top-5-data-science-trends-for-2020



Updated in 2020:  on  14 March 2020

7 June 2017, 2 September 2014


Tuesday, July 20, 2021

Deep Learning - Introduction and Bibliography



What is Deep Learning?


Deep learning is a form of machine learning for nonlinear high dimensional data reduction and prediction.

Using  Bayesian probabilistic perspective in deep learning provides a number of advantages. Specifically statistical interpretation and properties, more efficient algorithms for optimisation and
hyper-parameter tuning, and an explanation of predictive performance. 

Traditional high dimensional statistical techniques; principal component analysis (PCA), partial least squares (PLS), reduced rank regression (RRR), projection pursuit regression (PPR) are shallow learners.

Their deep learning counterparts exploit multiple layers of of data reduction which leads to performance gains. Stochastic gradient descent (SGD) training and optimisation and Dropout (DO) provides model and variable selection. Bayesian regularization is central to finding networks and provides a framework for optimal bias-variance trade-off to achieve good out-of sample performance.

To illustrate the use of bayesian perspective,  an analysis of first time international bookings on Airbnb. is presented in the paper.


https://arxiv.org/pdf/1706.00473.pdf



Deep Learning Introduction
___________________


___________________



How to get started with Deep Learning for Data Science?



-1. Learn Python and R ;)

0. Andrew Ng and Coursera

- https://lnkd.in/eUe9YZE

1. Siraj Raval: YouTube channel. Specifically this playlists:

- The Math of Intelligence: https://lnkd.in/eYPJbsW

- Intro to Deep Learning: https://lnkd.in/e4Sg9qy

2. François Chollet's book: Deep Learning with Python (and R soon):

- https://lnkd.in/gfV2ery
- https://lnkd.in/e6_YGqx

3. IBM Cognitive Class:

- https://lnkd.in/eNKPSnJ
- https://lnkd.in/eBVRf-R

4. Medium blogs:

- https://lnkd.in/eaUx5aN
- https://lnkd.in/eGaQwts

5. DataCamp:

- https://lnkd.in/eWVz7e5
- https://lnkd.in/ezXBq6M

Info collected from a Linkedin Post

https://www.linkedin.com/feed/update/urn:li:activity:6363784952114401280

--------------------------------------


Updated 21 July 2021,  2 February 2018
5 June 2017

Data Science - Online Study Programs, Notes and Video Courses - Free Also


UC Berkeley School of Information - Master of Information and Data Science (MIDS) - Curriculum

The online Master of Information and Data Science (MIDS) is designed to educate data science leaders
https://ischoolonline.berkeley.edu/data-science/curriculum/




https://towardsdatascience.com/functions-of-data-science-4afd5341a659

https://towardsdatascience.com/how-youtube-recommends-videos-b6e003a5ab2f

2018

You Need To Keep Learning In Data Science
https://datafloq.com/read/why-you-need-to-keep-learning-in-data-science/

10 Free Must-Read Books for Machine Learning and Data Science
April 2017
https://www.kdnuggets.com/2017/04/10-free-must-read-books-machine-learning-data-science.html 



2016
Learn R Free
https://www.datacamp.com/courses/free-introduction-to-r

http://tryr.codeschool.com/levels/1/challenges/2

Edureka YouTube Video

https://www.youtube.com/watch?v=TGo9F0QyBuE


Businesses Will Need One Million Data Scientists by 2018
International Data Corporation (IDC) predicts a need for 181,000 people with deep analytical skills in the US by 2018 and a requirement for five times that number of positions with data management and interpretation capabilities.
http://www.kdnuggets.com/2016/01/businesses-need-one-million-data-scientists-2018.html



Data analytics is  growing. Now computer applications in industry are broughtly divided into transaction application and intelligence applications. Business intelligence, data mining, data analytics, data science etc. are the subjects that are in the area of intelligence applications of computers in business organizations.
_______________

_______________



Updated  14 Feb 2016, 7 Feb 2016



NPTEL IIT Madras Course: Introduction to Data Analytics

http://nptel.ac.in/courses/110106064/







Harvard Stat 221 “Statistical Computing and Visualization”:  Online Lecture Links
http://harvarddatascience.com/2013/05/05/harvard-stat-221-statistical-computing-and-visualization-all-lectures-online/


Data Analysis
26 Resources 310+ Hours 24,298 Learners
Learn how to manipulate and analyze data better with this free online curriculum
https://www.springboard.com/learning-paths/data-analysis/





The Open Source Data Science Masters
Curriculum for Data Science
Follow me on Twitter @clarecorthell   - Follow the author of this blog on   @knoltweet

The Open-Source Data Science Masters
The open-source curriculum for learning Data Science.
Foundational in both theory and technologies, the OSDSM breaks down the core competencies necessary to make data useful.
http://datasciencemasters.org/

Updated  21 July 2021
13 July 2018, 2 February 2018
14 Apr 2016,  14 Feb 2016

Monday, May 31, 2021

Python Programming Language - Tutorials



The principal disadvantage of MATLAB against Python are the costs. Python is completely free, whereas MATLAB can be very expensive. Python is a very attractive alternative of MATLAB: Python is not only free of costs, but its code is open source. Python is continually becoming more powerful by a rapidly growing number of specialized modules.


Python tutorial: https://www.tutorialspoint.com/python/index.htm

--------------

http://www.python-course.eu/course.php


http://www.python-course.eu/python3_course.php


http://www.python-course.eu/advanced_topics.php

http://www.python-course.eu/numerical_programming.php

http://www.python-course.eu/python_tkinter.php


Ud 1 June 2021
Pub 12.6.2016

Sunday, October 4, 2020

Top IoT Systems and Components Vendors



IIoT Platforms Gartner

By 2025, 50% of industrial enterprises will use industrial Internet of Things (IIoT) platforms to improve factory operations, up from 10% in 2020.


Market Definition/Description
Gartner defines the IIoT platform market as a set of integrated software capabilities to improve asset management decision making within asset-intensive industries. IIoT platforms also provide operational visibility and control for plants, infrastructure and equipment.

IIoT Platforms
The IIoT platform  cost-effectively collects higher volumes of high-velocity, complex machine data from networked IoT endpoints. The IIoT platform also orchestrates historically siloed data sources to enable better accessibility, and improve insights and actions across a heterogeneous asset group through specialized analysis of the data.

The IIoT platform:
Monitors IoT endpoints and event streams
Analyzes data at the edge and in the cloud
Integrates and engages IT and OT systems in data sharing and consumption
Enables application development and deployment
Can enrich and supplement OT functions for improved asset management life cycle strategies and processes

The IIoT platform, in concert with the IoT edge and through enterprise IT/OT integration, prepares asset-intensive industries to become digital businesses. Digital capabilities are achieved by enhancing and connecting their core business with customers, suppliers and business partners.

The IIoT platform software that resides on and near devices — such as controllers, routers, access points, gateways and edge compute systems — is considered part of the “distributed IIoT platform.”

The platform provider must exhibit demonstrable value in terms of integration and interoperability with such applications, which include:

Enterprise asset management (EAM)
Computerized maintenance management systems (CMMSs)
Fleet management
Condition-based maintenance (CBM)
Manufacturing execution systems (MES)
Maintenance, repair and operations (MRO)
Product life cycle management (PLM)
Application portfolio management (APM)
Field service management (FSM)
Building management systems (BMSs)


IIoT Platform Capabilities
The IIoT platform  is composed of the following technology functions:

Device management — This function includes software that enables manual and automated tasks to create, provision, configure, troubleshoot and manage fleets of IoT devices and gateways remotely, in bulk or individually, and securely.

Integration — This function includes software, tools and technologies, such as communications protocols, APIs and application adapters, which minimally address the data, process, enterprise application and IIoT ecosystem integration requirements across cloud and on-premises implementations for end-to-end IIoT solutions. These IIoT solutions include IIoT devices (for example, communications modules and controllers), IIoT gateways, IIoT edge and IIoT platforms.

Data management — This function includes capabilities that support:
Ingesting IoT endpoint and edge device data
Storing data from edge to enterprise platforms
Providing data accessibility (by devices, IT and OT systems, and external parties, when required)
Tracking lineage and flow of data
Enforcing data and analytics governance policies to ensure the quality, security, privacy and currency of data

Analytics — This function includes processing of data streams, such as device, enterprise and contextual data, to provide insights into asset state by monitoring use, providing indicators, tracking patterns and optimizing asset use. A variety of techniques, such as rule engines, event stream processing, data visualization and machine learning, may be applied.

Application enablement and management — This function includes software that enables business applications in any deployment model to analyze data and accomplish IoT-related business functions. Core software components manage the OS, standard input and output or file systems to enable other software components of the platform. The application platform (for example, application platform as a service [aPaaS]) includes application-enabling infrastructure components, application development, runtime management and digital twins. The platform allows users to achieve “cloud scale” scalability and reliability and deploy and deliver IoT solutions quickly and seamlessly.
Security — This function includes the software, tools and practices facilitated to audit and ensure compliance. This function also establishes preventive, detective and corrective controls and actions to ensure privacy and the security of data across the IIoT solution.



2019

https://www.gartner.com/reviews/market/industrial-iot-platforms

Hitachi Again Named a “Visionary” in Gartner Magic Quadrant for IIoT Platforms 2019
https://www.hitachivantara.com/ext/gartner-magic-quadrant-for-industrial-iot.html

https://www.ptc.com/en/resources/iiot/white-paper/gartner-mq-for-iiot

--------------------------

IBM

Google

Intel

Microsoft

Cisco

Apple

SAP

Oracle

Samsung

Hewlett Packard

Ericson

Amazon.Com

GE

Qualcomm

AT&T

Orange

Blackberry

Facebook

Dell

Verizon

--------------------

News

February 2016

http://www.ecommercetimes.com/story/83088.html


IoT Players

http://electronicsofthings.com/category/industry-players/

http://internetofthingswiki.com/iot-companies-you-must-know/653



5 Oct 2020
25 March 2016


Monday, November 19, 2018

Big Data - Introduction




Big data usually includes data sets with sizes beyond the ability of commonly-used software tools to capture, curate, manage, and process the data within a tolerable elapsed time. Big data sizes are a constantly moving target, as of 2012 ranging from a few dozen terabytes to many petabytes of data in a single data set. With this difficulty, a new platform of "big data" tools has arisen to handle sensemaking over large quantities of data, as in the Apache Hadoop Big Data Platform.


In 2012, Gartner updated its definition as follows: "Big data are high-volume, high-velocity, and/or high-variety information assets that require new forms of processing to enable enhanced decision making, insight discovery and process optimization."

A 2016 definition states that "Big data represents the information assets characterized by such a high volume, velocity and variety to require specific technology and analytical methods for its transformation into value".

A 2018 definition states "Big data is where parallel computing tools are needed to handle data", and notes, "This represents a distinct and clearly defined change in the computer science used, via parallel programming theories, and losses of some of the guarantees and capabilities made by Codd’s relational model."

(Source:  http://en.wikipedia.org/wiki/Big_data  )

Big Data Repositories


Big data repositories have existed in many forms for year built by corporations for their use with a special need. Commercial vendors historically offered parallel database management systems for big data beginning in the 1990s.

Teradata Corporation in 1984 marketed the parallel processing DBC 1012 system. Teradata systems were the first to store and analyze 1 terabyte of data in 1992. Hard disk drives were 2.5 GB in 1991 so the definition of big data continuously evolves according to Kryder's Law. Teradata installed the first petabyte class RDBMS based system in 2007. As of 2017, there are a few dozen petabyte class Teradata relational databases installed, the largest of which exceeds 50 PB. Systems up until 2008 were 100% structured relational data. Since then, Teradata has added unstructured data types including XML, JSON, and Avro.

In 2000, Seisint Inc. (now LexisNexis Group) developed a C++-based distributed file-sharing framework for data storage and query. The system stores and distributes structured, semi-structured, and unstructured data across multiple servers. Users can build queries in a C++ dialect called ECL.  In 2004, LexisNexis acquired Seisint Inc. and in 2008 acquired ChoicePoint, Inc.and their high-speed parallel processing platform. The two platforms were merged into HPCC (or High-Performance Computing Cluster) Systems and in 2011, HPCC was open-sourced under the Apache v2.0 License. Quantcast File System was available about the same time.

CERN and other physics experiments have collected big data sets and they analyzed via high performance computing (supercomputers). But big data movement presently uses  the commodity map-reduce architectures.

In 2004, Google published a paper on a process called MapReduce. The MapReduce concept provides a parallel processing model  to process huge amounts of data. With MapReduce, queries are split and distributed across parallel nodes and processed in parallel (the Map step). The results are then gathered and delivered as the output (the Reduce step). An implementation of the MapReduce framework was adopted by an Apache open-source project named Hadoop. Apache Spark was developed in 2012 in response to limitations in the MapReduce paradigm, as it adds the ability to set up many operations (not just map followed by reduce).

MIKE2.0 is an open approach to information management that acknowledges the need for revisions due to big data implications identified in an article titled "Big Data Solution Offering". The methodology addresses handling big data in terms of useful permutations of data sources, complexity in interrelationships, and difficulty in deleting (or modifying) individual records.

https://en.wikipedia.org/wiki/Big_data


Big Data - Dimensions


Big data - Four dimensions: Volume, Velocity, Variety, and Veracity (IBM document)
Examples of big data in enterprises

Volume: Enterprises are awash with ever-growing data of all types, easily amassing terabytes—even petabytes—of information.

12 terabytes of Tweets created each day has to analysed to get improved product sentiment analysis
Convert 350 billion annual meter readings to better predict power consumption

Velocity: Sometimes 2 minutes is too late. For time-sensitive processes such as catching fraud, big data must be used as it streams into your enterprise in order to maximize its value.

Examples:
Scrutinize 5 million trade events created each day to identify potential fraud
Analyze 500 million daily call detail records in real-time to predict customer churn faster

Variety: Big data is any type of data - structured and unstructured data such as text, sensor data, audio, video, click streams, log files and more. New insights are found when analyzing these data types together.

Monitor 100’s of live video feeds from surveillance cameras to target points of interest
Exploit the 80% data growth in images, video and documents to improve customer satisfaction


Veracity:  Establishing trust in big data presents a huge challenge as the variety and number of sources grows.



McKinsey Article on Big Data
http://www.mckinsey.com/insights/mgi/research/technology_and_innovation/big_data_the_next_frontier_for_innovation

28.2.2013





11 Feb 2016

Evolution of Big

http://www.ibmbigdatahub.com/infographic/evolution-big-data


https://hbr.org/2013/12/analytics-30

Analytics 1.0—the era of “business intelligence.”

Analytics 1.0 started gaining an objective, deep understanding of important business phenomena and giving managers the fact-based comprehension to go beyond intuition when making decisions. For the first time, data about production processes, sales, customer interactions, and more were recorded, aggregated, and analyzed.


Updated  20 November 2018,  11 Feb 2016, 28 Feb 2013

Big Data - Analysis - Articles, Books and Research Papers - Bibliography


"Big data is where parallel computing tools are needed to handle data" - 2018 definition.

Big Data - Introduction

Big Data - Wikipedia Article



http://www.bigdata-madesimple.com/research-papers-that-changed-the-world-of-big-data/

It is a collection of research papers in the area of Big Data

MapReduce: Simplified Data Processing on Large Clusters

This paper presents MapReduce, a programming model and its implementation for large-scale distributed clusters. The main idea is to have a general execution model for codes that need to process a large amount of data over hundreds of machines.

The Google File System

It presents Google File System, a scalable distributed file system for large distributed data-intensive applications, which provides fault tolerance while running on inexpensive commodity hardware, and it delivers high aggregate performance to a large number of clients.

Bigtable: A Distributed Storage System for Structured Data

This paper presents the simple data model provided by Bigtable, which gives clients dynamic control over data layout and format, and the design and implementation of Bigtable.

Dynamo: Amazon’s Highly Available Key-value Store

This paper presents the design and implementation of Dynamo, a highly available key-value storage system that some of Amazon's core services use to provide an "always-on" experience.

The Chubby lock service for loosely-coupled distributed systems

Chubby is a distributed lock service; it does a lot of the hard parts of building distributed systems and provides its users with a familiar interface (writing files, taking a lock, file permissions). The paper describes it, focusing on the API rather than the implementation details.

Chukwa: A large-scale monitoring system

This paper describes the design and initial implementation of Chukwa, a data collection system for monitoring and analyzing large distributed systems. Chukwa is built on top of Hadoop, an open source distributed filesystem and MapReduce implementation, and inherits Hadoop’s scalability and robustness.

Cassandra - A Decentralized Structured Storage System

Cassandra is a distributed storage system for managing very large amounts of structured data spread out across many commodity servers, while providing highly available service with no single point of failure.

HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads

There are two schools of thought regarding what technology to use for data analysis. Proponents of parallel databases argue that the strong emphasis on performance and efficiency of parallel databases makes them well-suited to perform such analysis. On the other hand, others argue that MapReduce-based systems are better suited due to their superior scalability, fault tolerance, and flexibility to handle unstructured data. This paper explores the feasibility of building a hybrid system.

S4: Distributed Stream Computing Platform.

This paper outlines the S4 architecture in detail, describes various applications, including real-life deployments, to show that the S4 design is surprisingly flexible and lends itself to run in large clusters built with commodity hardware.

Dremel: Interactive Analysis of Web-Scale Datasets

This paper describes the architecture and implementation of Dremel, a scalable, interactive ad-hoc query system for analysis of read-only nested data, and explains how it complements MapReduce-based computing.

Large-scale Incremental Processing Using Distributed Transactions and Notifications

Percolator is a system for incrementally processing updates to a large data set, and deployed it to create the Google web search index. This indexing system based on incremental processing replaced Google's batch-based indexing system.

Pregel: A System for Large-Scale Graph Processing

This paper presents a computational model suitable to solve many practical computing problems that concerns large graphs.

Spanner: Google’s Globally-Distributed Database

It explains about Spanner, Google’s scalable, multi-version, globally-distributed, and synchronously-replicated database. It is the first system to distribute data at global scale and sup-port externally-consistent distributed transactions.

Shark: Fast Data Analysis Using Coarse-grained Distributed Memory

Shark is a research data analysis system built on a novel coarse-grained distributed shared-memory abstraction. Shark marries query processing with deep data analysis, providing a unified system for easy data manipulation using SQL and pushing sophisticated analysis closer to data.

The PageRank Citation Ranking: Bringing Order to the Web

This paper describes PageRank, a method for rating Web pages objectively and mechanically, effectively measuring the human interest and attention devoted to them.

A Few Useful Things to Know about Machine Learning

This paper summarizes twelve key lessons that machine learning researchers and practitioners have learned, which include pitfalls to avoid, important issues to focus on, and answers to common questions.

Random Forests

This paper describes a method of building a forest of uncorrelated trees using a CART like procedure, combined with randomized node optimization and bagging. In addition, it combines several ingredients, which form the basis of the modern practice of random forests.

A Relational Model of Data for Large Shared Data Banks

Written by EF Codd in 1970, this paper was a breakthrough in Relational Data Base systems. He was the man who first conceived of the relational model for database management.

Map-Reduce for Machine Learning on Multicore

The paper focuses on developing a general and exact technique for parallel programming of a large class of machine learning algorithms for multicore processors. The central idea is to allow a future programmer or user to speed up machine learning applications by "throwing more cores" at the problem rather than search for specialized optimizations.

Megastore: Providing Scalable, Highly Available Storage for Interactive Services

This paper describes Megastore, a storage system developed to blend the scalability of a NoSQL datastore with the convenience of a traditional RDBMS in a novel way.

Finding a needle in Haystack: Facebook’s photo storage

This paper describes Haystack, an object storage system optimized for Facebook’s Photos application. Facebook currently stores over 260 billion images, which translates to over 20 petabytes of data.

Spark: Cluster Computing with Working Sets

This paper focuses on applications that reuse a working set of data across multiple parallel operations and proposes a new framework called Spark that supports these applications while retaining the scalability and fault tolerance of MapReduce.

The Unified Logging Infrastructure for Data Analytics at Twitter

This paper presents Twitter’s production logging infrastructure and its evolution from application-specific logging to a unified “client events” log format, where messages are captured in common, well-formatted, flexible Thrift messages.

F1: A Distributed SQL Database That Scales

F1 is a distributed relational database system built at Google to support the AdWords business. F1 is a hybrid database that combines high availability, the scalability of NoSQL systems like Bigtable, and the consistency and usability of traditional SQL databases.

MLbase: A Distributed Machine-learning System

This paper presents MLbase, a novel system harnessing the power of machine learning for both end-users and ML researchers.

Scalable Progressive Analytics on Big Data in the Cloud

This paper presents a new approach that gives more control to data scientists to carefully choose from a huge variety of sampling strategies in a domain-specific manner.

Big data: The next frontier for innovation, competition, and productivity

This is paper one of the most referenced documents in the world of Big Data. It describes current and potential applications of Big Data.

The Promise and Peril of Big Data

This paper summarizes the insights of the Eighteenth Annual Roundtable on Information Technology, which sought to understand the implications of the emergence of  “Big Data” and new techniques of inferential analysis.

TDWI Checklist Report: Big Data Analytics

This paper provides six guidelines on implementing Big Data Analytics. It helps you take the first steps toward achieving a lasting competitive edge with analytics.

http://www.bigdata-madesimple.com/research-papers-that-changed-the-world-of-big-data/



Updated on 20 November 2018
Last updated 3 February 2015

Wednesday, June 7, 2017

Google BIGQUERY - Enterprise Cloud Data Warehouse



BigQuery is Google's fully managed, petabyte scale, low cost enterprise data warehouse for analytics. BigQuery is serverless. There is no infrastructure to manage and you don't need a database administrator, so you can focus on analyzing data to find meaningful insights using familiar SQL. BigQuery is a powerful Big Data analytics platform used by all types of organizations, from startups to Fortune 500 companies.

Speed & Scale

BigQuery can scan TB in seconds and PB in minutes. Load your data from Google Cloud Storage or Google Cloud Datastore, or stream it into BigQuery to enable real-time analysis of your data. With BigQuery you can easily scale your database from GBs to PBs.

https://cloud.google.com/bigquery/


Building and scaling new business models to gain insights from disparate data faster, while reducing IT costs, requires an architecture that can go from prototype to petabyte scale as your needs evolve. Google BigQuery’s serverless architecture can help ensure that your enterprise data warehouse withstands growth at any scale. Informatica helps you unlock the power of hybrid data with high performance, highly scalable data management solutions that efficiently move and manage large volumes of data to Google BigQuery. Informatica and Google BigQuery is the best combination for modernizing your data architecture.

https://www.brighttalk.com/webcast/10477/258043/modernize-your-data-architecture-with-google-bigquery-and-informatica

Friday, October 28, 2016

Network Analysis and Social Network Analysis - Bibliography and Videos




_________________


_________________




http://snap.stanford.edu/

http://snap.stanford.edu/class/cs224w-2015/projects.html

https://web.stanford.edu/class/cs224w/

https://web.stanford.edu/class/cs224w/intro_handout/intro_handout.pdf

http://historicalnetworkresearch.org/resources/first-steps/

Network Structure Inference, A Survey: Motivations, Methods, and
Applications
Ivan Brugere, University of Illinois at Chicago
Brian Gallagher, Lawrence Livermore National Laboratory
Tanya Y. Berger-Wolf, University of Illinois at Chicago
https://arxiv.org/pdf/1610.00782.pdf


Social Network Analysis: Methods and Applications

Stanley Wasserman, Katherine Faust
Cambridge University Press, 25-Nov-1994 - Social Science - 825 pages


Social network analysis, which focuses on relationships among social entities, is used widely in the social and behavioral sciences, as well as in economics, marketing, and industrial engineering. Social Network Analysis: Methods and Applications reviews and discusses methods for the analysis of social networks with a focus on applications of these methods to many substantive examples. As the first book to provide a comprehensive coverage of the methodology and applications of the field, this study is both a reference book and a textbook.
https://books.google.co.in/books?hl=en&lr=&id=CAm2DpIqRUIC

The structure and function of complex networks
M. E. J. Newman
Department of Physics, University of Michigan, Ann Arbor, MI 48109, U.S.A. and
Santa Fe Institute, 1399 Hyde Park Road, Santa Fe, NM 87501, U.S.A.



Analyzing Participation of Students in Online Courses Using Social
Network Analysis Techniques
Reihaneh Rabbany k., Mansoureh Takaffoli and Osmar R. Za¨ıane,
Department of Computing Science, University of Alberta, Canada
rabbanyk,takaffol,zaiane@ualberta.ca

Social Network Analysis and Mining to Support the
Assessment of On-line Student Participation


Big Data over Networks

Shuguang Cui, Alfred O. Hero, III, Zhi-Quan Luo
Cambridge University Press, 14-Jan-2016 - Computers - 457 pages


Utilising both key mathematical tools and state-of-the-art research results, this text explores the principles underpinning large-scale information processing over networks and examines the crucial interaction between big data and its associated communication, social and biological networks. Written by experts in the diverse fields of machine learning, optimisation, statistics, signal processing, networking, communications, sociology and biology, this book employs two complementary approaches: first analysing how the underlying network constrains the upper-layer of collaborative big data processing, and second, examining how big data processing may boost performance in various networks. Unifying the broad scope of the book is the rigorous mathematical treatment of the subjects, which is enriched by in-depth discussion of future directions and numerous open-ended problems that conclude each chapter. Readers will be able to master the fundamental principles for dealing with big data over large systems, making it essential reading for graduate students, scientific researchers and industry practitioners alike.


Networks, Crowds, and Markets: Reasoning about a Highly Connected World
Full Book
David Easley
Dept. of Economics
Cornell University

Jon Kleinberg
Dept. of Computer Science
Cornell University
Cambridge University Press, 2010
Draft version: June 10, 2010.

Fundamentals of Predictive Text Mining

Sholom M. Weiss, Nitin Indurkhya, Tong Zhang
Springer, 07-Sep-2015 - Computers - 239 pages


This successful textbook on predictive text mining offers a unified perspective on a rapidly evolving field, integrating topics spanning the varied disciplines of data science, machine learning, databases, and computational linguistics. Serving also as a practical guide, this unique book provides helpful advice illustrated by examples and case studies.

This highly anticipated second edition has been thoroughly revised and expanded with new material on deep learning, graph models, mining social media, errors and pitfalls in big data evaluation, Twitter sentiment analysis, and dependency parsing discussion. The fully updated content also features in-depth discussions on issues of document classification, information retrieval, clustering and organizing documents, information extraction, web-based data-sourcing, and prediction and evaluation.

Topics and features: presents a comprehensive, practical and easy-to-read introduction to text mining; includes chapter summaries, useful historical and bibliographic remarks, and classroom-tested exercises for each chapter; explores the application and utility of each method, as well as the optimum techniques for specific scenarios; provides several descriptive case studies that take readers from problem description to systems deployment in the real world; describes methods that rely on basic statistical techniques, thus allowing for relevance to all languages (not just English); contains links to free downloadable industrial-quality text-mining software and other supplementary instruction material.

Fundamentals of Predictive Text Mining is an essential resource for IT professionals and managers, as well as a key text for advanced undergraduate computer science students and beginning graduate students.

http://www.leonidzhukov.net/hse/2016/sna/


Advanced Database Marketing: Innovative Methodologies and Applications for Managing Customer Relationships

Koen W. De Bock
Routledge, Mar 23, 2016 - 348 pages
https://books.google.co.in/books?id=4hHPCwAAQBAJ



Tuesday, June 21, 2016

NoSQL Databases


The term was started in 2009

10 things you should know about NoSQL databases
By Guy Harrison, August 26, 2010
http://www.techrepublic.com/blog/10-things/10-things-you-should-know-about-nosql-databases/


Introduction - Presentation
http://martinfowler.com/articles/nosql-intro-original.pdf


2013  Introduction to NoSQL - Video Presentation by Martin Fowler

__________________

__________________

GOTO Conferences

http://martinfowler.com/nosql.html

https://www.mongodb.com/

http://www.nosqlweekly.com/  - Subscribe to their weekly. They provide a python weekly also.


Tuesday, May 10, 2016

Hadoop Adoption and Ecosystem


What is Hadoop?


“Hadoop” refers to the growing ecosystem of open source and commercial software platforms, tools, and maintenance available from the Apache Software Foundation (open source) and several software vendor firms.

Hadoop began its journey by proving its worth as a Spartan but highly scalable data platform for
reporting and analytics in Internet firms and other digital organizations. The journey is now taking
Hadoop into a wider range of industries,

 Hadoop adoption is accelerating.  60% of users surveyed will have Hadoop in production by 2016.

“The high-end relational databases are really expensive in configurations big enough to deal with big data. Data warehouse appliances are almost as expensive. We all need a more economical platform, which is the main reason we’re all considering Hadoop.”

 Hadoop regularly appears as a complementary extension of a data warehouse when warehouse data that doesn’t necessarily require the warehouse is migrated to Hadoop.

Hadoop usage is proliferating across enterprises.

 A growing number of users rely on Hadoop to spur enterprise business and technology innovations.


Benefits of Hadoop



Advanced analytics.

Hadoop supports advanced analytics, based on techniques for data mining,
statistics, complex SQL etc. This includes the exploratory analytics with big data.
 It also includes related disciplines such as
information exploration and discovery and data visualization.

Data warehousing and integration.

Many users feel Hadoop complements a data warehouse well, is
a big data source for analytics , and is a computational platform for transforming data .

Data Scalability.

With Hadoop, users feel they can capture more data than in the past . This
is the perception, whether use cases involve analytics, warehouses, or active data archiving .
Technology and economics intersect because users feel they can achieve extreme scalability
while running on low-cost hardware and software .

New and exotic data types.

Hadoop helps organizations get business value from data that is new to them or previously unmanageable, simply because Hadoop supports widely diverse data and file types. In particular, Hadoop is adept with schema-free data staging  and machine data from
robots, sensors, meters, and other devices .

Business applications.

Hadoop contributes to a number of business applications and activities,
including sentiment analytics , understanding consumer behavior via clickstreams ,
more numerous business insights,  recognition of sales and market opportunities , fraud
detection , and greater ROI for big data .



Important points being summarized from

Hadoop for theEnterprise: Making Data Management Massively Scalable, Agile, Feature-Rich, and Cost-Effective


By Philip Russom

http://www.sas.com/content/dam/SAS/en_us/doc/whitepaper2/tdwi-hadoop-for-enterprise-107708.pdf

updated 10 May 2016

Friday, March 6, 2015

Big Data Application in Medical Practice




Design and Development of a Medical Big Data Processing System Based on Hadoop

Qin Yao, Yu Tian, Peng-Fei Li, Li-Li Tian, Yang-Ming Qian, Jing-Song Li

Journal of Medical Systems (Impact Factor: 1.37). 03/2015; 39(3):220. DOI: 10.1007/s10916-015-0220-8

Tuesday, September 2, 2014

Cloudera Big Data Products and Services



Cloudera Express - Freedownload Big Data Manager

Cloudera Express
The Best Way to Get Started with Hadoop
Cloudera Express is a free download that combines CDH, Cloudera’s 100% open source and enterprise-ready distribution of Apache Hadoop with Cloudera Manager, which provides robust cluster management capabilities like automated deployment, centralized administration, monitoring, and diagnostic tools.

Cloudera Express gives you everything you need to get started with Hadoop. With Cloudera Express, you get a fully capable platform that’s optimized to help you demonstrate the value of the technology.

Whether you are evaluating Hadoop to accelerate data processing, optimize the performance of your data warehouse, or perform new types of analysis on data sets that were previously out of reach, Cloudera Express is your key to successfully deploying Hadoop to solve your first use cases.
http://www.cloudera.com/content/cloudera/en/products-and-services/cloudera-express.html


Cloudera Enterprise
Hadoop for the Enterprise
Cloudera Enterprise helps you become information-driven by leveraging the best of the open source community with the enterprise capabilities you need to succeed with Apache Hadoop in your organization. Designed specifically for mission-critical environments, Cloudera Enterprise includes CDH, the world’s most popular open source Hadoop-based platform, as well as advanced system management and data management tools plus dedicated support and community advocacy from our world-class team of Hadoop developers and experts. Cloudera is your partner on the path to big data.
http://www.cloudera.com/content/cloudera/en/products-and-services/cloudera-enterprise.html


Cloudera Live (beta)
Try a live demo of Hadoop, right now.
Cloudera Live is a new way to get started with Apache Hadoop, online. No downloads, no installations, no waiting. Watch tutorial videos and work with real-world examples of the complete Hadoop stack included with CDH, Cloudera’s completely open source Hadoop platform, to:
Learn Hue, the Hadoop User Interface developed by Cloudera
Query data using popular projects like Apache Hive, Apache Pig, Impala, Apache Solr, and Apache Spark (new!)
Develop workflows using Apache Oozie
http://www.cloudera.com/content/cloudera/en/products-and-services/cloudera-live.html


Cloudera Blog
http://vision.cloudera.com/

Cloudera's Stategy to compete in Big Data space
http://www.forbes.com/sites/danwoods/2014/05/09/clouderas-strategy-for-conquering-big-data-the-enterprise/


Designing a Scalable and Agile Big Data Platform
December 2011
http://www.citoresearch.com/data-science/designing-scalable-and-agile-big-data-platform