Big Data Research in Italy: A Perspective
EngineeringPublished 1 June 2016Open access
Sonia Bergamaschi, Emanuele Carlini, Michelangelo Ceci, Barbara Furletti, Fosca Giannotti, Donato Malerba
Citations22
SJR quartileQ1
SJR score1.77
SNIP2.52
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A sample of distinct applications that address the issue of managing huge amounts of data in Italy, collected in relation to diverse domains are offered.
Abstract
The aim of this article is to synthetically describe the research projects that a selection of Italian universities \nis undertaking in the context of big data. Far from being exhaustive, this article has the objective \nof offering a sample of distinct applications that address the issue of managing huge amounts of data in \nItaly, collected in relation to diverse domains.
Keywords
Decision Sciences
International Journal on Semantic Web and Information SystemsLinked Data - The Story So Far
4,562 Citations2009Christian Bizer, Tom Heath +1 more
The authors describe progress to date in publishing Linked Data on the Web, review applications that have been developed to exploit the Web of Data, and map out a research agenda for the Linked data community as it moves forward.
Spark: cluster computing with working sets
4,236 Citations2010Matei Zaharia, Mosharaf Chowdhury +3 more
Spark can outperform Hadoop by 10x in iterative machine learning jobs, and can be used to interactively query a 39 GB dataset with sub-second response time.
Communications of the ACMThe KDD process for extracting useful knowledge from volumes of data
1,899 Citations1996Usama M. Fayyad, Gregory Piatetsky-Shapiro +1 more
A new generation of computational techniques and tools is required to support the extraction of useful knowledge from the rapidly growing volumes of data, the subject of the emerging field of knowledge discovery in databases (KDD) and data mining.
Human mobility, social ties, and link prediction
674 Citations2011Dashun Wang, Dino Pedreschi +3 more
It is shown that mobility measures alone yield surprising predictive power, comparable to traditional network-based measures, and the prediction accuracy can be significantly improved by learning a supervised classifier based on combined mobility and network measures.
Proceedings of the VLDB EndowmentWebTables
634 Citations2008Michael Cafarella, Alon Halevy +3 more
The WEBTABLES system develops new techniques for keyword search over a corpus of tables, and shows that they can achieve substantially higher relevance than solutions based on a traditional search engine.
Nature CommunicationsReturners and explorers dichotomy in human mobility
487 Citations2015Luca Pappalardo, Filippo Simini +4 more
It is shown that returners and explorers play a distinct quantifiable role in spreading phenomena and that a correlation exists between their mobility patterns and social interactions.
Predicting solar generation from weather forecasts using machine learning
471 Citations2011Navin Sharma, Pranshu Sharma +2 more
This paper explores automatically creating site-specific prediction models for solar power generation from National Weather Service weather forecasts using machine learning techniques, and shows that SVM-based prediction models built using seven distinct weather forecast metrics are 27% more accurate for the authors' site than existing forecast-based models.
Web-scale Data Integration: You can only afford to Pay As You Go
301 Citations2007Jayant Madhavan, Shawn R. Jeffery +5 more
This paper proposes a new data integration architecture, PAYGO, which is inspired by the concept of dataspaces and emphasizes pay-as-you-go data management as means for achieving web-scale data integration.
The Knowledge Engineering ReviewA multidisciplinary survey on discrimination analysis
296 Citations2013Andrea Romei, Salvatore Ruggieri
This survey is to provide a guidance and a glue for researchers and anti-discrimination data analysts on concepts, problems, application areas, datasets, methods, and approaches from a multidisciplinary perspective.
Progress in Photovoltaics Research and ApplicationsSolar and photovoltaic forecasting through post‐processing of the Global Environmental Multiscale numerical weather prediction model
241 Citations2011Sophie Pelland, George Galanis +1 more
IEEE Transactions on Power SystemsEntropy and Correntropy Against Minimum Square Error in Offline and Online Three-Day Ahead Wind Power Forecasting
199 Citations2009Ricardo J. Bessa, Vladimiro Miranda +1 more
Renyi's entropy is combined with a Parzen windows estimation of the error pdf to form the basis of two criteria (minimum entropy and maximum correntropy) under which neural networks are trained.
Movement Data Anonymity through Generalization
139 Citations2010Anna Monreale, Gennady Andrienko +5 more
Journal of Database ManagementFrom Data Quality to Big Data Quality
132 Citations2015Carlo Batini, Anisa Rula +2 more
The nature of the relationship between Data Quality and several research coordinates that are relevant in Big Data, such as the variety of data types, data sources and application domains, are examined, focusing on maps, semi-structured texts, linked open data, sensor & sensor networks and official statistics.
IEEE Transactions on Knowledge and Data EngineeringA Blocking Framework for Entity Resolution in Highly Heterogeneous Information Spaces
132 Citations2013George Papadakis, Ekaterini Ioannou +3 more
This paper systemize blocking methods for clean-clean ER (an inherently quadratic task) over highly heterogeneous information spaces (HHIS) through a novel framework that consists of two orthogonal layers: the effectiveness layer encompasses methods for building overlapping blocks with small likelihood of missed matches and the efficiency layer comprises a rich variety of techniques that significantly restrict the required number of pairwise comparisons.
IEEE Transactions on Knowledge and Data EngineeringMeta-Blocking: Taking Entity Resolutionto the Next Level
128 Citations2013George Papadakis, Georgia Koutrika +2 more
This paper introduces meta-blocking as a generic procedure that intervenes between the creation and the processing of blocks, transforming an initial set of blocks into a new one with substantially fewer comparisons and equally high effectiveness.
IEEE Systems JournalPrivacy-Preserving Mining of Association Rules From Outsourced Transaction Databases
124 Citations2012Fosca Giannotti, Laks V. S. Lakshmanan +3 more
This paper proposes an attack model based on background knowledge and devise a scheme for privacy preserving outsourced mining, which ensures that each transformed item is indistinguishable with respect to the attacker's background knowledge, from at least k-1 other transformed items.
Synthesis lectures on data managementBig Data Integration
124 Citations2013Dong Xin, Divesh Srivastava
Using big data to study the link between human mobility and socio-economic development
107 Citations2015Luca Pappalardo, Dino Pedreschi +2 more
Information Technology and LibrariesUsability Studies of Faceted Browsing: A Literature Review
102 Citations2010Jody Condit Fagan
The article proposes methodological considerations for practicing librarians and provides examples of goals, tasks, and measurements for user studies of faceted browsing in library catalogs.
Movement data anonymity through generalization
100 Citations2009Gennady Andrienko, Natalia Andrienko +3 more
This position paper briefly presents an approach for the generalization of movement data that can be adopted for obtaining k-anonymity in spatio-temporal datasets and can be used to realize a framework for publishing of spatio/temporal data while preserving privacy.
Data Mining and Knowledge DiscoveryDiscrimination- and privacy-aware patterns
96 Citations2014Sara Hajian, Josep Domingo‐Ferrer +3 more
It is argued that privacy and discrimination risks should be tackled together, and a methodology for doing so while publishing frequent pattern mining results is presented, and pattern sanitization methods based on $$k$$k-anonymity yield both privacy- and discrimination-protected patterns, while introducing reasonable (controlled) pattern distortion.
BMC Public HealthChronic disease prevalence from Italian administrative databases in the VALORE project: a validation through comparison of population estimates with general practice databases and national survey
95 Citations2013Rosa Gini, Paolo Francesconi +14 more
This study supports the use of data from Italian administrative databases to estimate geographic differences in population prevalence of ischaemic heart disease, treated diabetes, diabetes mellitus and heart failure and the algorithm for COPD requires further refinement.
Data Science and EngineeringOn the Meaningfulness of “Big Data Quality” (Invited Paper)
80 Citations2015Donatella Firmani, Massimo Mecella +2 more
The overall aim of the paper is to identify further research directions in the area of big data quality, by providing at the same time an up-to-date state of the art on data quality.
EPJ Data SciencePrivacy-by-design in big data analytics and social mining
75 Citations2014Anna Monreale, Salvatore Rinzivillo +3 more
The privacy-by-design paradigm is proposed to develop technological frameworks for countering the threats of undesirable, unlawful effects of privacy violation, without obstructing the knowledge discovery opportunities of social mining and big data analytical technologies.
Lecture notes in geoinformation and cartographyGeographic Information Science at the Heart of Europe
74 Citations2013Danny Vandenbroucke, Bénédicte Bucher +1 more
This chapter introduces the concept of micromapping, the acquisition of geometrically correct geometric data of small geographic entities and presents MAPIT, a method for micro-mapping with smartphones with high geometric precision.
Semantic services, interoperability, and web applications emerging concepts
68 Citations2011Amit Sheth
Telecommunications PolicyDiscovering urban and country dynamics from mobile phone data with spatial correlation patterns
64 Citations2014Roberto Trasarti, Ana‐Maria Olteanu‐Raimond +6 more
An analytical process aimed at extracting interconnections between different areas of the city that emerge from highly correlated temporal variations of population local densities based on spatiotemporal aggregations of presence is proposed.
2014 International Conference on Computing, Networking and Communications (ICNC)Big Data as an e-Health Service
60 Citations2014W. Liu, E. K. Park
This paper explains why the existing Big Data technologies such as Hadoop, MapReduce, STORM and the like cannot be simply applied to e-Health services directly, and describes the additional capabilities as required in order to make Big Data services for e- health become practical.
Proceedings of the VLDB EndowmentSupervised meta-blocking
56 Citations2014George Papadakis, George Papastefanatos +1 more
It is shown that supervised meta-blocking can achieve high performance with small training sets that can be manually created and is compared with baseline and competitor methods over 10 large-scale datasets, both real and synthetic.
Data Mining and Knowledge DiscoveryNetwork regression with predictive clustering trees
52 Citations2012Daniela Stojanova, Michelangelo Ceci +2 more
This paper proposes a data mining algorithm, called NCLUS, that explicitly considers autocorrelation when building regression models from network data, based on the concept of predictive clustering trees (PCTs) that can be used for clustering, prediction and multi-target prediction, including multi- target regression andMulti-target classification.
Challenge: Processing web texts for classifying job offers
46 Citations2015Flora Amato, Roberto Boselli +6 more
This paper applies and compares several techniques, namely explicit-rules, machine learning, and LDA-based algorithms to classify a real dataset of Web job offers collected from 12 heterogeneous sources against a standard classification system of occupations.
Information Processing & ManagementA model-based evaluation of data quality activities in KDD
36 Citations2014Mario Mezzanzanica, Roberto Boselli +2 more
The experimental outcomes show the effectiveness of the model-based approach for data quality as they provide a fine-grained analysis of both the source dataset and the cleansing procedures, enabling domain experts to identify the most relevant quality issues as well as the action points for improving the cleansing activities.
IEEE Transactions on Industrial InformaticsWind Power Forecasting in a Residential Location as Part of the Energy Box Management Decision Tool
32 Citations2014Christos S. Ioakimidis, Luís J. Oliveira +1 more
This work employs a multilayer feed-forward back-propagation neural network for classification that utilizes the global forecast system predictions on wind speed and direction to identify patterns of the wind behavior at the location considered in order to obtain a stochastic distribution of the daily wind speed.
City users' classification with mobile phone data
30 Citations2015Lorenzo Gabrielli, Barbara Furletti +3 more
Lecture notes in geoinformation and cartographyPrivacy-Preserving Distributed Movement Data Aggregation
28 Citations2013Anna Monreale, Hui Wang +5 more
This work proposes a novel approach to privacy-preserving analytical processing within a distributed setting, and tackles the problem of obtaining aggregated information about vehicle traffic in a city from movement data collected by individual vehicles and shipped to a central server.
Big Data Techniques For Supporting Accurate Predictions of Energy Production From Renewable Sources
25 Citations2014Michelangelo Ceci, Roberto Corizzo +7 more
This paper uses HBase over Hadoop framework on a cluster of commodity servers in order to provide a system that can be used as a basis for running machine learning algorithms, and performs one-day ahead forecast of PV energy production based on Artificial Neural Networks.
Proceedings of the International Conference on Automated Planning and SchedulingPlanning meets Data Cleansing
24 Citations2014Roberto Boselli, Mirko Cesarini +2 more
The concept of cost-optimal Universal Cleanser — a collection of cleansing actions for each data inconsistency — is formalised as a planning problem and a motivating government application in which it has been used is presented.
Analysis of GSM calls data for understanding user mobility behavior
23 Citations2013Barbara Furletti, Lorenzo Gabrielli +2 more
This paper proposes a strategy for mobility behavior identification based on aggregated calling profiles of mobile phone users, and shows how these call profiles permit to design a two step process based on a bootstrap phase and a running phase for classifying users into behavior categories.
Lecture notes in computer scienceCitizen in Sensor Networks
18 Citations2013Jordi Nin, Daniel Villatoro
The system processes multi-modal data in real-time, performing paralleled task recognition and modality synchronization, showing high performance recognizing subjects, objects, and interactions, and its reliability to be applied in real case scenarios.
Lecture notes in computer scienceTransportation Planning Based on GSM Traces: A Case Study on Ivory Coast
17 Citations2013Mirco Nanni, Roberto Trasarti +6 more
An analysis process that exploits mobile phone transaction (trajectory) data to infer a transport demand model for the territory under monitoring and shows a case study on Ivory Coast, with emphasis on its major urbanization Abidjan.
CINECA IRIS Institutial research information system (University of Pisa)Using t-closeness anonymity to control for non-discrimination
17 Citations2015Salvatore Ruggieri
It is shown that t-closeness implies bdf(t)-protection, for a bound function bdf() depending on the discrimination measure f() at hand, to adapt inference control methods, such as the Mondrian multidimensional generalization technique and the Sabre bucketization and redistribution framework to the purpose of non-discrimination data protection.
Data Quality Sensitivity Analysis on Aggregate Indicators
16 Citations2012Mario Mezzanzanica, Roberto Boselli +2 more
This paper proposes a methodology exploiting Finite State Systems to quantitatively estimate how computed variables and indicators might be affected by the uncertainty related to low data quality, independently from the data cleansing methodology used.
Expert Systems with ApplicationsDiscrimination discovery in scientific project evaluation: A case study
13 Citations2013Andrea Romei, Salvatore Ruggieri +1 more
This paper contributes by presenting a case study on gender discrimination in a dataset of scientific research proposals, and by distilling from the case study a general discrimination discovery process, thus obtaining the best of the two worlds.
Transactions on data privacyUsing t-closeness anonymity to control for non-discrimination
10 Citations2014RuggieriSalvatore
Innovative power operating center management exploiting big data techniques
9 Citations2014Michelangelo Ceci, Nunziato Cassavia +6 more
The Vi-POC (Virtual Power Operating Center) project is described, a project conceived to assist energy producers and decision makers in the energy market and how it faces with challenges posed by the specific application.
Studies in classification, data analysis, and knowledge organizationEnhancing Big Data Exploration with Faceted Browsing
4 Citations2018Sonia Bergamaschi, Giovanni Simonini +1 more
This article presents a new approach to visualize huge amount of data, based on a Bayesian suggestion algorithm and the widely used enterprise search platform Solr, and demonstrates how the proposed Bayesian suggest algorithm became a key ingredient in a big data scenario, where generally a query can generate so many results that the user can be confused.
Discovering the topics of a data source: A statistical approach?
1 Citations2014Sonia Bergamaschi, Davide Ferrari +2 more
A novel data-driven technique based on composite likelihood to estimate the weights and other main features of the graphs is proposed, making the resulting approach less sensitive to overfitting.
