Contact The DL Team Contact Us | Switch to tabbed view

top of pageABSTRACT

It is our great pleasure to welcome you to SIGMOD 2016, the 2016 edition of the ACM SIGMOD International Conference on Management of Data, at San Francisco, the heart of Silicon Valley. A unique bridge, cable cars, a sparkling bay, a famous prison, and steep streets with beautiful Victorian houses are just a few of the things that make San Francisco one of the world's greatest cities to visit. Located in the Bay Area along the Northern California coast, many attractions beyond those in the city can be easily reached: the world-famous wine country, including Napa and Sonoma Valley, beautiful beaches, the world's largest trees and wonderful hikes in nearby National Parks.

The conference itself is held in the Hyatt Regency Hotel, located in the financial district in downtown San Francisco. The hotel is very conveniently located nearby to many activities and landmarks, making this the ideal spot not only for an exciting conference but also for exploring the area. The banquet will be held in the California Academy of Sciences. The California Academy of Sciences is a Planetarium, Aquarium and Natural History Museum under a Living Roof in San Francisco, and it is among the largest museums of natural history in the world, housing over 26 million specimens.

This year's technical program features 137 papers in 22 sessions; 116 research papers, 17 industrial papers, and 4 invited papers. Research and industrial papers will be mixed together instead of separated into separate tracks as in past years. All papers will be presented at conference sessions, as well as at one of two plenary poster events. There are also 31 demos and 10 tutorials. Instead of running these bconcurrently with technical presentations, demos are now a separate plenary event in the conference and tutorials have been moved before and after the conference. There is just one panel, the Careers in Industry panel. Finally, this year's conference will feature one keynote, from the inimitable Jeff Dean (Google, USA), who will talk about how deep learning and software systems can support each other, with particular thoughts on the implications for data management. The net effect of these changes is that there are more plenary events, and at most three concurrent events ongoing (from 6 or more concurrent tracks in previous conferences). We hope this will make the SIGMOD experience more communal.

Other technical events of interest include the SIGMOD programming contest, featuring a $7,000 prize for the winning team plus a $3,000 dollar prize for the runner-up donated by Microsoft, with contest participants presenting their work during the poster session; the undergraduate poster competition, featuring 12 undergraduate posters from an outstanding group of young database researchers; the new research symposium on Tuesday night; and the business meeting, also on Tuesday evening.

We made a number of changes in the way we assembled this year's technical program with an eye towards improving review and paper quality. As in past years we had two submission deadlines, but we spread them out more (one in July and one in November). Each submission included one round of revisions, during which reviewers could request specific changes to papers. The vast majority of accepted papers underwent revision, which we believe significantly improved the quality of papers. All submissions were also discussed in a series of conference calls organized by our group leaders; this forced reviewers to defend their positions and helped achieve consensus. A post-review survey revealed 70% of reviewers felt this improved the review process, and 80% reported they worked harder in preparing and defending their reviews. We also asked reviewers and group leaders to request additional reviews when an additional opinion was needed; many papers received four reviews, and some papers received as many as six. Finally, we introduced an online conference organization tool, Confer, which we used to get the community's input about how papers should be grouped together, and which attendees can use throughout the conference to navigate the program. SIGMOD 2016 is preceded and followed by 10 workshops, including the PhD symposium on Sunday, June 26. These workshops provide an opportunity for small groups of like-minded researchers to discuss areas of interest and new ideas, and we are pleased to see a diverse ecosystem of workshops with competitive submissions evolving at SIGMOD.

top of pageSOURCE MATERIALS

FRONT MATTER
PDFPDF  (Title Page, Copyright, Welcome Message from the SIGMOD Chairs, contents, Organization, Sponsors)
BACK MATTER
PDFPDF  (Author Index)
APPEARS IN
Digital Content

top of pageAUTHORS

General Chairs


Author image not provided  Fatma Özcan

No contact information provided yet.

Bibliometrics: publication history
Publication years1995-2016
Publication count40
Citation Count482
Available for download24
Downloads (6 Weeks)418
Downloads (12 Months)1,783
Downloads (cumulative)12,225
Average downloads per article509.38
Average citations per article12.05
View colleagues of Fatma Özcan


Author image not provided  Georgia Koutrika

No contact information provided yet.

Bibliometrics: publication history
Publication years2012-2014
Publication count3
Citation Count6
Available for download1
Downloads (6 Weeks)4
Downloads (12 Months)6
Downloads (cumulative)39
Average downloads per article39.00
Average citations per article2.00
View colleagues of Georgia Koutrika
Program Chairs


Author image not provided  Sam Madden

No contact information provided yet.

Bibliometrics: publication history
Publication years2001-2016
Publication count153
Citation Count7,052
Available for download119
Downloads (6 Weeks)1,849
Downloads (12 Months)15,689
Downloads (cumulative)193,325
Average downloads per article1,624.58
Average citations per article46.09
View colleagues of Sam Madden

top of pageREFERENCES

References are not available

top of pageCITED BY

Citings are not available

top of pageINDEX TERMS

Index Terms are not available

top of pagePUBLICATION

Title SIGMOD/PODS'16 International Conference on Management of Data
San Francisco, CA, USA — June 26 - July 01, 2016
Pages2272
Sponsor SIGMOD ACM Special Interest Group on Management of Data
PublisherACM New York, NY, USA
ISBN 978-1-4503-3531-7
Conference MODInternational Conference on Management of Data MOD logo
Overall Acceptance Rate 1,104 of 5,662 submissions, 19%
Year Submitted Accepted Rate
SIGMOD '96 290 47 16%
SIGMOD '97 202 42 21%
SIGMOD '00 248 42 17%
SIGMOD '01 293 44 15%
SIGMOD '02 240 42 18%
SIGMOD '03 342 53 15%
SIGMOD '06 446 58 13%
SIGMOD '07 480 70 15%
SIGMOD '08 435 78 18%
SIGMOD '09 430 118 27%
SIGMOD '10 384 80 21%
SIGMOD '11 375 93 25%
SIGMOD '12 289 48 17%
SIGMOD '13 372 76 20%
SIGMOD '14 421 107 25%
SIGMOD '15 415 106 26%
Overall 5,662 1,104 19%

APPEARS IN
Digital Content

top of pageREVIEWS


Reviews are not available for this item
Computing Reviews logo

top of pageCOMMENTS

Be the first to comment To Post a comment please sign in or create a free Web account

top of pageTable of Contents

Proceedings of the 2016 International Conference on Management of Data
Table of Contents
SESSION: Keynote - Jeff Dean
Building Machine Learning Systems that Understand
Jeff Dean
Pages: 1-1
doi>10.1145/2882903.2932259
Full text: PDFPDF

Over the past five years, deep learning and large-scale neural networks have made significant advances in speech recognition, computer vision, language understanding and translation, robotics, and many other fields. Deep learning allows the use of very ...
expand
SESSION: Session 1 - Scalable Analytics and Machine Learning
Learning Linear Regression Models over Factorized Joins
Maximilian Schleich, Dan Olteanu, Radu Ciucanu
Pages: 3-18
doi>10.1145/2882903.2882939
Full text: PDFPDF

We investigate the problem of building least squares regression models over training datasets defined by arbitrary join queries on database tables. Our key observation is that joins entail a high degree of redundancy in both computation and data representation, ...
expand
To Join or Not to Join?: Thinking Twice about Joins before Feature Selection
Arun Kumar, Jeffrey Naughton, Jignesh M. Patel, Xiaojin Zhu
Pages: 19-34
doi>10.1145/2882903.2882952
Full text: PDFPDF

Closer integration of machine learning (ML) with data processing is a booming area in both the data management industry and academia. Almost all ML toolkits assume that the input is a single table, but many datasets are not stored as single tables due ...
expand
Real-time Video Recommendation Exploration
Yanxiang Huang, Bin Cui, Jie Jiang, Kunqian Hong, Wenyu Zhang, Yiran Xie
Pages: 35-46
doi>10.1145/2882903.2903743
Full text: PDFPDF

Video recommendation has attracted growing attention in recent years. However, conventional techniques have limitations in real-time processing, accuracy or scalability for the large-scale video data. To address the deficiencies of current recommendation ...
expand
Towards Globally Optimal Crowdsourcing Quality Management: The Uniform Worker Setting
Akash Das Sarma, Aditya Parameswaran, Jennifer Widom
Pages: 47-62
doi>10.1145/2882903.2882953
Full text: PDFPDF

We study crowdsourcing quality management, that is, given worker responses to a set of tasks, our goal is to jointly estimate the true answers for the tasks, as well as the quality of the workers. Prior work on this problem relies primarily on applying ...
expand
Building the Enterprise Fabric for Big Data with Vertica and Spark Integration
Jeff LeFevre, Rui Liu, Cornelio Inigo, Lupita Paz, Edward Ma, Malu Castellanos, Meichun Hsu
Pages: 63-75
doi>10.1145/2882903.2903744
Full text: PDFPDF

Enterprise customers increasingly require greater flexibility in the way they access and process their Big Data while at the same time they continue to request advanced analytics and access to diverse data sources. Yet customers also still require the ...
expand
Truss Decomposition of Probabilistic Graphs: Semantics and Algorithms
Xin Huang, Wei Lu, Laks V.S. Lakshmanan
Pages: 77-90
doi>10.1145/2882903.2882913
Full text: PDFPDF

A key operation in network analysis is the discovery of cohesive subgraphs. The notion of $k$-truss has gained considerable popularity in this regard, based on its rich structure and efficient computability. However, many complex networks such as social, ...
expand
Efficient and Progressive Group Steiner Tree Search
Rong-Hua Li, Lu Qin, Jeffrey Xu Yu, Rui Mao
Pages: 91-106
doi>10.1145/2882903.2915217
Full text: PDFPDF

The Group Steiner Tree (GST) problem is a fundamental problem in database area that has been successfully applied to keyword search in relational databases and team search in social networks. The state-of-the-art algorithm for the GST problem is a parameterized ...
expand
SESSION: Session 2 - Privacy and Security
Publishing Attributed Social Graphs with Formal Privacy Guarantees
Zach Jorgensen, Ting Yu, Graham Cormode
Pages: 107-122
doi>10.1145/2882903.2915215
Full text: PDFPDF

Many data analysis tasks rely on the abstraction of a graph to represent relations between entities, with attributes on the nodes and edges. Since the relationships encoded are often sensitive, we seek effective ways to release representative graphs ...
expand
Publishing Graph Degree Distribution with Node Differential Privacy
Wei-Yen Day, Ninghui Li, Min Lyu
Pages: 123-138
doi>10.1145/2882903.2926745
Full text: PDFPDF

Graph data publishing under node-differential privacy (node-DP) is challenging due to the huge sensitivity of queries. However, since a node in graph data oftentimes represents a person, node-DP is necessary to achieve personal data protection. In this ...
expand
Principled Evaluation of Differentially Private Algorithms using DPBench
Michael Hay, Ashwin Machanavajjhala, Gerome Miklau, Yan Chen, Dan Zhang
Pages: 139-154
doi>10.1145/2882903.2882931
Full text: PDFPDF

Differential privacy has become the dominant standard in the research community for strong privacy protection. There has been a flood of research into query answering algorithms that meet this standard. Algorithms are becoming increasingly complex, and ...
expand
PrivTree: A Differentially Private Algorithm for Hierarchical Decompositions
Jun Zhang, Xiaokui Xiao, Xing Xie
Pages: 155-170
doi>10.1145/2882903.2882928
Full text: PDFPDF

Given a set D of tuples defined on a domain Omega, we study differentially private algorithms for constructing a histogram over Omega to approximate the tuple distribution in D. Existing solutions for the problem mostly adopt a hierarchical decomposition ...
expand
Adaptive Indexing over Encrypted Numeric Data
Panagiotis Karras, Artyom Nikitin, Muhammad Saad, Rudrika Bhatt, Denis Antyukhov, Stratos Idreos
Pages: 171-183
doi>10.1145/2882903.2882932
Full text: PDFPDF

Today, outsourcing query processing tasks to remote cloud servers becomes a viable option; such outsourcing calls for encrypting data stored at the server so as to render it secure against eavesdropping adversaries and/or an honest-but-curious server ...
expand
Practical Private Range Search Revisited
Ioannis Demertzis, Stavros Papadopoulos, Odysseas Papapetrou, Antonios Deligiannakis, Minos Garofalakis
Pages: 185-198
doi>10.1145/2882903.2882911
Full text: PDFPDF

We consider a data owner that outsources its dataset to an untrusted server. The owner wishes to enable the server to answer range queries on a single attribute, without compromising the privacy of the data and the queries. There are several schemes ...
expand
Privacy Preserving Subgraph Matching on Large Graphs in Cloud
Zhao Chang, Lei Zou, Feifei Li
Pages: 199-213
doi>10.1145/2882903.2882956
Full text: PDFPDF

The wide presence of large graph data and the increasing popularity of storing data in the cloud drive the needs for graph query processing on a remote cloud. But a fundamental challenge is to process user queries without compromising sensitive information. ...
expand
SESSION: Session 3 - Logical and Physical Database Design
The Snowflake Elastic Data Warehouse
Benoit Dageville, Thierry Cruanes, Marcin Zukowski, Vadim Antonov, Artin Avanes, Jon Bock, Jonathan Claybaugh, Daniel Engovatov, Martin Hentschel, Jiansheng Huang, Allison W. Lee, Ashish Motivala, Abdul Q. Munir, Steven Pelley, Peter Povinec, Greg Rahn, Spyridon Triantafyllis, Philipp Unterbrunner
Pages: 215-226
doi>10.1145/2882903.2903741
Full text: PDFPDF

We live in the golden age of distributed computing. Public cloud platforms now offer virtually unlimited compute and storage resources on demand. At the same time, the Software-as-a-Service (SaaS) model brings enterprise-class systems to users who previously ...
expand
Closing the functional and Performance Gap between SQL and NoSQL
Zhen Hua Liu, Beda Hammerschmidt, Doug McMahon, Ying Liu, Hui Joe Chang
Pages: 227-238
doi>10.1145/2882903.2903731
Full text: PDFPDF

Oracle release 12cR1 supports JSON data management that enables users to store, index and query JSON data along with relational data. The integration of the JSON data model into the RDBMS allows a new paradigm of data management where data is storable, ...
expand
Have Your Data and Query It Too: From Key-Value Caching to Big Data Management
Dipti Borkar, Ravi Mayuram, Gerald Sangudi, Michael Carey
Pages: 239-251
doi>10.1145/2882903.2904443
Full text: PDFPDF

Couchbase Server is a rethinking of the database given the current set of realities. Memory today is much cheaper than disks were when traditional databases were designed back in the 1970's, and networks are much faster and much more reliable than ever ...
expand
Ambry: LinkedIn's Scalable Geo-Distributed Object Store
Shadi A. Noghabi, Sriram Subramanian, Priyesh Narayanan, Sivabalan Narayanan, Gopalakrishna Holla, Mammad Zadeh, Tianwei Li, Indranil Gupta, Roy H. Campbell
Pages: 253-265
doi>10.1145/2882903.2903738
Full text: PDFPDF

The infrastructure beneath a worldwide social network has to continually serve billions of variable-sized media objects such as photos, videos, and audio clips. These objects must be stored and served with low latency and high throughput by a system ...
expand
SQL Schema Design: Foundations, Normal Forms, and Normalization
Henning Köhler, Sebastian Link
Pages: 267-279
doi>10.1145/2882903.2915239
Full text: PDFPDF

Normalization helps us find a database schema at design time that can process the most frequent updates efficiently at run time. Unfortunately, relational normalization only works for idealized database instances in which duplicates and null markers ...
expand
SQLShare: Results from a Multi-Year SQL-as-a-Service Experiment
Shrainik Jain, Dominik Moritz, Daniel Halperin, Bill Howe, Ed Lazowska
Pages: 281-293
doi>10.1145/2882903.2882957
Full text: PDFPDF

We analyze the workload from a multi-year deployment of a database-as-a-service platform targeting scientists and data scientists with minimal database experience. Our hypothesis was that relatively minor changes to the way databases are delivered can ...
expand
Automatic Generation of Normalized Relational Schemas from Nested Key-Value Data
Michael DiScala, Daniel J. Abadi
Pages: 295-310
doi>10.1145/2882903.2882924
Full text: PDFPDF

Self-describing key-value data formats such as JSON are becoming increasingly popular as application developers choose to avoid the rigidity imposed by the relational model. Database systems designed for these self-describing formats, such as MongoDB, ...
expand
SESSION: Session 4 - New Storage and Network Architectures
Data Blocks: Hybrid OLTP and OLAP on Compressed Storage using both Vectorization and Compilation
Harald Lang, Tobias Mühlbauer, Florian Funke, Peter A. Boncz, Thomas Neumann, Alfons Kemper
Pages: 311-326
doi>10.1145/2882903.2882925
Full text: PDFPDF

This work aims at reducing the main-memory footprint in high performance hybrid OLTP & OLAP databases, while retaining high query performance and transactional throughput. For this purpose, an innovative compressed columnar storage format for cold ...
expand
GeckoFTL: Scalable Flash Translation Techniques For Very Large Flash Devices
Niv Dayan, Philippe Bonnet, Stratos Idreos
Pages: 327-342
doi>10.1145/2882903.2915219
Full text: PDFPDF

The volume of metadata needed by a flash translation layer (FTL) is proportional to the storage capacity of a flash device. Ideally, this metadata should reside in the device's integrated RAM to enable fast access. However, as flash devices scale to ...
expand
SHARE Interface in Flash Storage for Relational and NoSQL Databases
Gihwan Oh, Chiyoung Seo, Ravi Mayuram, Yang-Suk Kee, Sang-Won Lee
Pages: 343-354
doi>10.1145/2882903.2882910
Full text: PDFPDF

Database consistency and recoverability require guaranteeing write atomicity for one or more pages. However, contemporary database systems consider write operations non-atomic. Thus, many database storage engines have traditionally relied on either journaling ...
expand
Accelerating Relational Databases by Leveraging Remote Memory and RDMA
Feng Li, Sudipto Das, Manoj Syamala, Vivek R. Narasayya
Pages: 355-370
doi>10.1145/2882903.2882949
Full text: PDFPDF

Memory is a crucial resource in relational databases (RDBMSs). When there is insufficient memory, RDBMSs are forced to use slower media such as SSDs or HDDs, which can significantly degrade workload performance. Cloud database services are deployed in ...
expand
FPTree: A Hybrid SCM-DRAM Persistent and Concurrent B-Tree for Storage Class Memory
Ismail Oukid, Johan Lasperas, Anisoara Nica, Thomas Willhalm, Wolfgang Lehner
Pages: 371-386
doi>10.1145/2882903.2915251
Full text: PDFPDF

The advent of Storage Class Memory (SCM) is driving a rethink of storage systems towards a single-level architecture where memory and storage are merged. In this context, several works have investigated how to design persistent trees in SCM as a fundamental ...
expand
Micro-architectural Analysis of In-memory OLTP
Utku Sirin, Pinar Tözün, Danica Porobic, Anastasia Ailamaki
Pages: 387-402
doi>10.1145/2882903.2882916
Full text: PDFPDF

Micro-architectural behavior of traditional disk-based online transaction processing (OLTP) systems has been investigated extensively over thepast couple of decades. Results show that traditional OLTP mostly under-utilize the available micro-architectural ...
expand
SESSION: Session 5 - Graphs 1: Infrastructure and Processing on Modern Hardware
iBFS: Concurrent Breadth-First Search on GPUs
Hang Liu, H. Howie Huang, Yang Hu
Pages: 403-416
doi>10.1145/2882903.2882959
Full text: PDFPDF

Breadth-First Search (BFS) is a key graph algorithm with many important applications. In this work, we focus on a special class of graph traversal algorithm - concurrent BFS - where multiple breadth-first traversals are performed simultaneously on the ...
expand
Tornado: A System For Real-Time Iterative Analysis Over Evolving Data
Xiaogang Shi, Bin Cui, Yingxia Shao, Yunhai Tong
Pages: 417-430
doi>10.1145/2882903.2882950
Full text: PDFPDF

There is an increasing demand for real-time iterative analysis over evolving data. In this paper, we propose a novel execution model to obtain timely results at given instants. We notice that a loop starting from a good initial guess usually converges ...
expand
EmptyHeaded: A Relational Engine for Graph Processing
Christopher R. Aberger, Susan Tu, Kunle Olukotun, Christopher Ré
Pages: 431-446
doi>10.1145/2882903.2915213
Full text: PDFPDF

There are two types of high-performance graph processing engines: low- and high-level engines. Low-level engines (Galois, PowerGraph, Snap) provide optimized data structures and computation models but require users to write low-level imperative code, ...
expand
GTS: A Fast and Scalable Graph Processing Method based on Streaming Topology to GPUs
Min-Soo Kim, Kyuhyeon An, Himchan Park, Hyunseok Seo, Jinwook Kim
Pages: 447-461
doi>10.1145/2882903.2915204
Full text: PDFPDF

A fast and scalable graph processing method becomes increasingly important as graphs become popular in a wide range of applications and their sizes are growing rapidly. Most of distributed graph processing methods require a lot of machines equipped with ...
expand
Graph Analytics Through Fine-Grained Parallelism
Zechao Shang, Feifei Li, Jeffrey Xu Yu, Zhiwei Zhang, Hong Cheng
Pages: 463-478
doi>10.1145/2882903.2915238
Full text: PDFPDF

Large graphs are getting increasingly popular and even indispensable in many applications, for example, in social media data, large networks, and knowledge bases. Efficient graph analytics thus becomes an important subject of study. To increase efficiency ...
expand
Hybrid Pulling/Pushing for I/O-Efficient Distributed and Iterative Graph Computing
Zhigang Wang, Yu Gu, Yubin Bao, Ge Yu, Jeffrey Xu Yu
Pages: 479-494
doi>10.1145/2882903.2882938
Full text: PDFPDF

Billion-node graphs are rapidly growing in size in many applications such as online social networks. Most graph algorithms generate a large number of messages during iterative computations. Vertex-centric distributed systems usually store graph data ...
expand
SESSION: Session 6 - Streaming 1: Systems and Outlier Detection
Scalable Pattern Sharing on Event Streams*
Medhabi Ray, Chuan Lei, Elke A. Rundensteiner
Pages: 495-510
doi>10.1145/2882903.2882947
Full text: PDFPDF

Complex Event Processing (CEP) has emerged as a technology of choice for high performance event analytics in time-critical decision-making applications. Yet it is becoming increasingly difficult to support high-performance event processing due to the ...
expand
How to Win a Hot Dog Eating Contest: Distributed Incremental View Maintenance with Batch Updates
Milos Nikolic, Mohammad Dashti, Christoph Koch
Pages: 511-526
doi>10.1145/2882903.2915246
Full text: PDFPDF

In the quest for valuable information, modern big data applications continuously monitor streams of data. These applications demand low latency stream processing even when faced with high volume and velocity of incoming changes and the user's desire ...
expand
Sharing-Aware Outlier Analytics over High-Volume Data Streams
Lei Cao, Jiayuan Wang, Elke A. Rundensteiner
Pages: 527-540
doi>10.1145/2882903.2882920
Full text: PDFPDF

Real-time analytics of anomalous phenomena on streaming data typically relies on processing a large variety of continuous outlier detection requests, each configured with different parameter settings. The processing of such complex outlier analytics ...
expand
THEMIS: Fairness in Federated Stream Processing under Overload
Evangelia Kalyvianaki, Marco Fiscato, Theodoros Salonidis, Peter Pietzuch
Pages: 541-553
doi>10.1145/2882903.2882943
Full text: PDFPDF

Federated stream processing systems, which utilise nodes from multiple independent domains, can be found increasingly in multi-provider cloud deployments, internet-of-things systems, collaborative sensing applications and large-scale grid systems. To ...
expand
SABER: Window-Based Hybrid Stream Processing for Heterogeneous Architectures
Alexandros Koliousis, Matthias Weidlich, Raul Castro Fernandez, Alexander L. Wolf, Paolo Costa, Peter Pietzuch
Pages: 555-569
doi>10.1145/2882903.2882906
Full text: PDFPDF

Modern servers have become heterogeneous, often combining multi-core CPUs with many-core GPGPUs. Such heterogeneous architectures have the potential to improve the performance of data-intensive stream processing applications, but they are not supported ...
expand
Range Thresholding on Streams
Miao Qiao, Junhao Gan, Yufei Tao
Pages: 571-582
doi>10.1145/2882903.2915965
Full text: PDFPDF

This paper studies a type of continuous queries called range thresholding on streams (RTS). Imagine the stream as an unbounded sequence of elements each of which is a real value. A query registers an interval, and must be notified as soon as a ...
expand
SESSION: Session 7 - Approximate Query Processing
Bridging the Archipelago between Row-Stores and Column-Stores for Hybrid Workloads
Joy Arulraj, Andrew Pavlo, Prashanth Menon
Pages: 583-598
doi>10.1145/2882903.2915231
Full text: PDFPDF

Data-intensive applications seek to obtain trill insights in real-time by analyzing a combination of historical data sets alongside recently collected data. This means that to support such hybrid workloads, database management systems (DBMSs) need to ...
expand
An Effective Syntax for Bounded Relational Queries
Yang Cao, Wenfei Fan
Pages: 599-614
doi>10.1145/2882903.2882942
Full text: PDFPDF

A query Q is boundedly evaluable under a set A of access constraints if for all datasets D that satisfy A, there exists a fraction DQ of D such that Q(D) = Q(DQ), and the size of DQ and time for identifying DQ ...
expand
Best Paper Wander Join: Online Aggregation via Random Walks
Feifei Li, Bin Wu, Ke Yi, Zhuoyue Zhao
Pages: 615-629
doi>10.1145/2882903.2915235
Full text: PDFPDF

Joins are expensive, and online aggregation over joins was proposed to mitigate the cost, which offers users a nice and flexible tradeoff between query efficiency and accuracy in a continuous, online fashion. However, the state-of-the-art approach, in ...
expand
Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters
Srikanth Kandula, Anil Shanbhag, Aleksandar Vitorovic, Matthaios Olma, Robert Grandl, Surajit Chaudhuri, Bolin Ding
Pages: 631-646
doi>10.1145/2882903.2882940
Full text: PDFPDF

We present a system that approximates the answer to complex ad-hoc queries in big-data clusters by injecting samplers on-the-fly and without requiring pre-existing samples. Improvements can be substantial when big-data queries take multiple passes over ...
expand
A Study of Sorting Algorithms on Approximate Memory
Shuang Chen, Shunning Jiang, Bingsheng He, Xueyan Tang
Pages: 647-662
doi>10.1145/2882903.2882908
Full text: PDFPDF

Hardware evolution has been one of the driving factors for the redesign of database systems. Recently, approximate storage emerges in the area of computer architecture. It trades off precision for better performance and/or energy consumption. Previous ...
expand
Distributed Wavelet Thresholding for Maximum Error Metrics
Ioannis Mytilinis, Dimitrios Tsoumakos, Nectarios Koziris
Pages: 663-677
doi>10.1145/2882903.2915230
Full text: PDFPDF

Modern data analytics involve simple and complex computations over enormous numbers of data records. The volume of data and the increasingly stringent response-time requirements place increasing emphasis on the efficiency of approximate query processing. ...
expand
Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee
Bolin Ding, Silu Huang, Surajit Chaudhuri, Kaushik Chakrabarti, Chi Wang
Pages: 679-694
doi>10.1145/2882903.2915249
Full text: PDFPDF

Data volumes are growing exponentially for our decision-support systems making it challenging to ensure interactive response time for ad-hoc queries without increasing cost of hardware. Aggregation queries with Group By that produce an aggregate value ...
expand
SESSION: Session 8 - Networks and the Web
Stop-and-Stare: Optimal Sampling Algorithms for Viral Marketing in Billion-scale Networks
Hung T. Nguyen, My T. Thai, Thang N. Dinh
Pages: 695-710
doi>10.1145/2882903.2915207
Full text: PDFPDF

Influence Maximization (IM), that seeks a small set of key users who spread the influence widely into the network, is a core problem in multiple domains. It finds applications in viral marketing, epidemic control, and assessing cascading failures within ...
expand
Spheres of Influence for More Effective Viral Marketing
Yasir Mehmood, Francesco Bonchi, David García-Soriano
Pages: 711-726
doi>10.1145/2882903.2915250
Full text: PDFPDF

What is the set of nodes of a social network that, under a probabilistic contagion model, would get infected if a given node $s$ gets infected? We call this set the sphere of influence of s. Due to the stochastic nature of the contagion model ...
expand
Continuous Influence Maximization: What Discounts Should We Offer to Social Network Users?
Yu Yang, Xiangbo Mao, Jian Pei, Xiaofei He
Pages: 727-741
doi>10.1145/2882903.2882961
Full text: PDFPDF

Imagine we are introducing a new product through a social network, where we know for each user in the network the purchase probability curve with respect to discount. Then, what discount should we offer to those social network users so that the adoption ...
expand
Holistic Influence Maximization: Combining Scalability and Efficiency with Opinion-Aware Models
Sainyam Galhotra, Akhil Arora, Shourya Roy
Pages: 743-758
doi>10.1145/2882903.2882929
Full text: PDFPDF

The steady growth of graph data from social networks has resulted in wide-spread research in finding solutions to the influence maximization problem. In this paper, we propose a holistic solution to the influence maximization (IM) problem. (1) We introduce ...
expand
Potential and Pitfalls of Domain-Specific Information Extraction at Web Scale
Astrid Rheinländer, Mario Lehmann, Anja Kunkel, Jörg Meier, Ulf Leser
Pages: 759-771
doi>10.1145/2882903.2903736
Full text: PDFPDF

In many domains, a plethora of textual information is available on the web as news reports, blog posts, community portals, etc. Information extraction (IE) is the default technique to turn unstructured text into structured fact databases, but systematically ...
expand
Robust and Noise Resistant Wrapper Induction
Tim Furche, Jinsong Guo, Sebastian Maneth, Christian Schallhart
Pages: 773-784
doi>10.1145/2882903.2915214
Full text: PDFPDF

Wrapper induction is the problem of automatically inferring a query from annotated web pages of the same template. This query should not only select the annotated content accurately but also other content following the same template. Beyond accurately ...
expand
SESSION: Session 9 - Data Discovery and Extraction
Goods: Organizing Google's Datasets
Alon Halevy, Flip Korn, Natalya F. Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy, Steven Euijong Whang
Pages: 795-806
doi>10.1145/2882903.2903730
Full text: PDFPDF

Enterprises increasingly rely on structured datasets to run their businesses. These datasets take a variety of forms, such as structured files, databases, spreadsheets, or even services that provide access to the data. The datasets often reside in different ...
expand
Multi-Source Uncertain Entity Resolution at Yad Vashem: Transforming Holocaust Victim Reports into People
Tomer Sagi, Avigdor Gal, Omer Barkol, Ruth Bergman, Alexander Avram
Pages: 807-819
doi>10.1145/2882903.2903737
Full text: PDFPDF

In this work we describe an entity resolution project performed at Yad Vashem, the central repository of Holocaust-era information. The Yad Vashem dataset is unique with respect to classic entity resolution, by virtue of being both massively multi-source ...
expand
A Hybrid Approach to Functional Dependency Discovery
Thorsten Papenbrock, Felix Naumann
Pages: 821-833
doi>10.1145/2882903.2915203
Full text: PDFPDF

Functional dependencies are structural metadata that can be used for schema normalization, data integration, data cleansing, and many other data management tasks. Despite their importance, the functional dependencies of a specific dataset are usually ...
expand
Ontological Pathfinding
Yang Chen, Sean Goldberg, Daisy Zhe Wang, Soumitra Siddharth Johri
Pages: 835-846
doi>10.1145/2882903.2882954
Full text: PDFPDF

Recent years have seen a drastic rise in the construction of web-scale knowledge bases (e.g., Freebase, YAGO, DBPedia). These knowledge bases store structured information about real-world people, places, organizations, etc. However, due to limitations ...
expand
Extracting Databases from Dark Data with DeepDive
Ce Zhang, Jaeho Shin, Christopher Ré, Michael Cafarella, Feng Niu
Pages: 847-859
doi>10.1145/2882903.2904442
Full text: PDFPDF

DeepDive is a system for extracting relational databases from dark data: the mass of text, tables, and images that are widely collected and stored but which cannot be exploited by standard relational tools. If the information in dark data --- scientific ...
expand
Estimating the Impact of Unknown Unknowns on Aggregate Query Results
Yeounoh Chung, Michael Lind Mortensen, Carsten Binnig, Tim Kraska
Pages: 861-876
doi>10.1145/2882903.2882909
Full text: PDFPDF

It is common practice for data scientists to acquire and integrate disparate data sources to achieve higher quality results. But even with a perfectly cleaned and merged data set, two fundamental questions remain: (1) is the integrated data set complete ...
expand
SESSION: Session 10 - Data Integration / Cleaning
Constraint-Variance Tolerant Data Repairing
Shaoxu Song, Han Zhu, Jianmin Wang
Pages: 877-892
doi>10.1145/2882903.2882955
Full text: PDFPDF

Integrity constraints, guiding the cleaning of dirty data, are often found to be imprecise as well. Existing studies consider the inaccurate constraints that are oversimplified, and thus refine the constraints via inserting more predicates (attributes). ...
expand
Interactive and Deterministic Data Cleaning
Jian He, Enzo Veltri, Donatello Santoro, Guoliang Li, Giansalvatore Mecca, Paolo Papotti, Nan Tang
Pages: 893-907
doi>10.1145/2882903.2915242
Full text: PDFPDF

We present Falcon, an interactive, deterministic, and declarative data cleaning system, which uses SQL update queries as the language to repair data. Falcon does not rely on the existence of a set of pre-defined data quality rules. On the contrary, it ...
expand
Sequential Data Cleaning: A Statistical Approach
Aoqian Zhang, Shaoxu Song, Jianmin Wang
Pages: 909-924
doi>10.1145/2882903.2915233
Full text: PDFPDF

Errors are prevalent in data sequences, such as GPS trajectories or sensor readings. Existing methods on cleaning sequential data employ a constraint on value changing speeds and perform constraint-based repairing. While such speed constraints are effective ...
expand
Learning-Based Cleansing for Indoor RFID Data
Asif Iqbal Baba, Manfred Jaeger, Hua Lu, Torben Bach Pedersen, Wei-Shinn Ku, Xike Xie
Pages: 925-936
doi>10.1145/2882903.2882907
Full text: PDFPDF

RFID is widely used for object tracking in indoor environments, e.g., airport baggage tracking. Analyzing RFID data offers insight into the underlying tracking systems as well as the associated business processes. However, the inherent uncertainty in ...
expand
PrivateClean: Data Cleaning and Differential Privacy
Sanjay Krishnan, Jiannan Wang, Michael J. Franklin, Ken Goldberg, Tim Kraska
Pages: 937-951
doi>10.1145/2882903.2915248
Full text: PDFPDF

Recent advances in differential privacy make it possible to guarantee user privacy while preserving the main characteristics of the data. However, most differential privacy mechanisms assume that the underlying dataset is clean. This paper explores the ...
expand
RDFind: Scalable Conditional Inclusion Dependency Discovery in RDF Datasets
Sebastian Kruse, Anja Jentzsch, Thorsten Papenbrock, Zoi Kaoudi, Jorge-Arnulfo Quiané-Ruiz, Felix Naumann
Pages: 953-967
doi>10.1145/2882903.2915206
Full text: PDFPDF

Inclusion dependencies (INDs) form an important integrity constraint on relational databases, supporting data management tasks, such as join path discovery and query optimization. Conditional inclusion dependencies (CINDs), which define including and ...
expand
Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach
Chengliang Chai, Guoliang Li, Jian Li, Dong Deng, Jianhua Feng
Pages: 969-984
doi>10.1145/2882903.2915252
Full text: PDFPDF

Crowdsourced entity resolution has recently attracted significant attentions because it can harness the wisdom of crowd to improve the quality of entity resolution. However existing techniques either cannot achieve high quality or incur huge monetary ...
expand
SESSION: Session 11 - Spatio / Temporal Databases
Topic Exploration in Spatio-Temporal Document Collections
Kaiqi Zhao, Lisi Chen, Gao Cong
Pages: 985-998
doi>10.1145/2882903.2882921
Full text: PDFPDF

Huge amounts of data with both spatial and temporal information (e.g., geo-tagged tweets) are being generated, and are often used to share and spread personal updates, spontaneous ideas, and breaking news. We refer to such data as spatio-temporal documents. ...
expand
ParTime: Parallel Temporal Aggregation
Markus Pilman, Martin Kaufmann, Florian Köhl, Donald Kossmann, Damien Profeta
Pages: 999-1010
doi>10.1145/2882903.2903732
Full text: PDFPDF

This paper presents ParTime, a parallel algorithm for temporal aggregation. Temporal aggregation is one of the most important, yet most complex temporal query operators. It has been extensively studied in the past, but so far there has only been one ...
expand
Data Polygamy: The Many-Many Relationships among Urban Spatio-Temporal Data Sets
Fernando Chirigati, Harish Doraiswamy, Theodoros Damoulas, Juliana Freire
Pages: 1011-1025
doi>10.1145/2882903.2915245
Full text: PDFPDF

The increasing ability to collect data from urban environments, coupled with a push towards openness by governments, has resulted in the availability of numerous spatio-temporal data sets covering diverse aspects of a city. Discovering relationships ...
expand
Distributed Evaluation of Top-k Temporal Joins
Julien Pilourdault, Vincent Leroy, Sihem Amer-Yahia
Pages: 1027-1039
doi>10.1145/2882903.2882912
Full text: PDFPDF

We study a particular kind of join, coined Ranked Temporal Join (RTJ), featuring predicates that compare time intervals and a scoring function associated with each predicate to quantify how well it is satisfied. RTJ queries are prevalent in a variety ...
expand
AT-GIS: Highly Parallel Spatial Query Processing with Associative Transducers
Peter Ogden, David Thomas, Peter Pietzuch
Pages: 1041-1054
doi>10.1145/2882903.2882962
Full text: PDFPDF

Users in many domains, including urban planning, transportation, and environmental science want to execute analytical queries over continuously updated spatial datasets. Current solutions for large-scale spatial query processing either rely on extensions ...
expand
Towards Best Region Search for Data Exploration
Kaiyu Feng, Gao Cong, Sourav S. Bhowmick, Wen-Chih Peng, Chunyan Miao
Pages: 1055-1070
doi>10.1145/2882903.2882960
Full text: PDFPDF

The increasing popularity and growth of mobile devices and location-based services enable us to utilize large-scale geo-tagged data to support novel location-based applications. This paper introduces a novel problem called the best region search ...
expand
Simba: Efficient In-Memory Spatial Analytics
Dong Xie, Feifei Li, Bin Yao, Gefei Li, Liang Zhou, Minyi Guo
Pages: 1071-1085
doi>10.1145/2882903.2915237
Full text: PDFPDF

Large spatial data becomes ubiquitous. As a result, it is critical to provide fast, scalable, and high-throughput spatial queries and analytics for numerous applications in location-based services (LBS). Traditional spatial databases and spatial analytics ...
expand
SESSION: Session 12 - Distributed Data Processing
Realtime Data Processing at Facebook
Guoqiang Jerry Chen, Janet L. Wiener, Shridhar Iyer, Anshul Jaiswal, Ran Lei, Nikhil Simha, Wei Wang, Kevin Wilfong, Tim Williamson, Serhat Yilmaz
Pages: 1087-1098
doi>10.1145/2882903.2904441
Full text: PDFPDF

Realtime data processing powers many use cases at Facebook, including realtime reporting of the aggregated, anonymized voice of Facebook users, analytics for mobile applications, and insights for Facebook page administrators. Many companies have developed ...
expand
SparkR: Scaling R Programs with Spark
Shivaram Venkataraman, Zongheng Yang, Davies Liu, Eric Liang, Hossein Falaki, Xiangrui Meng, Reynold Xin, Ali Ghodsi, Michael Franklin, Ion Stoica, Matei Zaharia
Pages: 1099-1104
doi>10.1145/2882903.2903740
Full text: PDFPDF

R is a popular statistical programming language with a number of extensions that support data processing and machine learning tasks. However, interactive data analysis in R is usually limited as the R runtime is single threaded and can only process data ...
expand
VectorH: Taking SQL-on-Hadoop to the Next Level
Andrei Costea, Adrian Ionescu, Bogdan Răducanu, Michał Switakowski, Cristian Bârca, Juliusz Sompolski, Alicja Łuszczak, Michał Szafrański, Giel de Nijs, Peter Boncz
Pages: 1105-1117
doi>10.1145/2882903.2903742
Full text: PDFPDF

Actian Vector in Hadoop (VectorH for short) is a new SQL-on-Hadoop system built on top of the fast Vectorwise analytical database system. VectorH achieves fault tolerance and storage scalability by relying on HDFS, and extends the state-of-the-art in ...
expand
Adaptive Logging: Optimizing Logging and Recovery Costs in Distributed In-memory Databases
Chang Yao, Divyakant Agrawal, Gang Chen, Beng Chin Ooi, Sai Wu
Pages: 1119-1134
doi>10.1145/2882903.2915208
Full text: PDFPDF

By maintaining the data in main memory, in-memory databases dramatically reduce the I/O cost of transaction processing. However, for recovery purposes, in-memory systems still need to flush the log to disk, which incurs a substantial number of I/Os. ...
expand
Big Data Analytics with Datalog Queries on Spark
Alexander Shkapsky, Mohan Yang, Matteo Interlandi, Hsuan Chiu, Tyson Condie, Carlo Zaniolo
Pages: 1135-1149
doi>10.1145/2882903.2915229
Full text: PDFPDF

There is great interest in exploiting the opportunity provided by cloud computing platforms for large-scale analytics. Among these platforms, Apache Spark is growing in popularity for machine learning and graph analytics. Developing efficient complex ...
expand
An Efficient MapReduce Cube Algorithm for Varied DataDistributions
Tova Milo, Eyal Altshuler
Pages: 1151-1165
doi>10.1145/2882903.2882922
Full text: PDFPDF

Data cubes allow users to discover insights from their data and are commonly used in data analysis. While very useful, the data cube is expensive to compute, in particular when the input relation is very large. To address this problem, we consider cube ...
expand
SESSION: Session 13 - Graphs 2: Subgraph-based Optimization Techniques
Diversified Top-k Subgraph Querying in a Large Graph
Zhengwei Yang, Ada Wai-Chee Fu, Ruifeng Liu
Pages: 1167-1182
doi>10.1145/2882903.2915216
Full text: PDFPDF

Subgraph querying in a large data graph is interesting for different applications. A recent study shows that top-k diversified results are useful since the number of matching subgraphs can be very large. In this work, we study the problem of top-k diversified ...
expand
Graph Indexing for Shortest-Path Finding over Dynamic Sub-Graphs
Mohamed S. Hassan, Walid G. Aref, Ahmed M. Aly
Pages: 1183-1197
doi>10.1145/2882903.2882933
Full text: PDFPDF

A variety of applications spanning various domains, e.g., social networks, transportation, and bioinformatics, have graphs as first-class citizens. These applications share a vital operation, namely, finding the shortest path between two nodes. In many ...
expand
Efficient Subgraph Matching by Postponing Cartesian Products
Fei Bi, Lijun Chang, Xuemin Lin, Lu Qin, Wenjie Zhang
Pages: 1199-1214
doi>10.1145/2882903.2915236
Full text: PDFPDF

In this paper, we study the problem of subgraph matching that extracts all subgraph isomorphic embeddings of a query graph q in a large data graph G. The existing algorithms for subgraph matching follow Ullmann's backtracking approach; that is, iteratively ...
expand
Adding Counting Quantifiers to Graph Patterns
Wenfei Fan, Yinghui Wu, Jingbo Xu
Pages: 1215-1230
doi>10.1145/2882903.2882937
Full text: PDFPDF

This paper proposes quantified graph patterns (QGPs), an extension of graph patterns by supporting simple counting quantifiers on edges. We show that QGPs naturally express universal and existential quantification, numeric and ratio aggregates, as well ...
expand
DUALSIM: Parallel Subgraph Enumeration in a Massive Graph on a Single Machine
Hyeonji Kim, Juneyoung Lee, Sourav S. Bhowmick, Wook-Shin Han, JeongHoon Lee, Seongyun Ko, Moath H.A. Jarrah
Pages: 1231-1245
doi>10.1145/2882903.2915209
Full text: PDFPDF

Subgraph enumeration is important for many applications such as subgraph frequencies, network motif discovery, graphlet kernel computation, and studying the evolution of social networks. Most earlier work on subgraph enumeration assumes that graphs are ...
expand
Distributed Set Reachability
Sairam Gurajada, Martin Theobald
Pages: 1247-1261
doi>10.1145/2882903.2915226
Full text: PDFPDF

In this paper, we focus on the efficient and scalable processing of set-reachability queries over a distributed, directed data graph. A "set-reachability query" is a generalized form of a reachability query, in which we consider two sets S and T of source ...
expand
SESSION: Session 14 - Main Memory Analytics
Fast Multi-Column Sorting in Main-Memory Column-Stores
Wenjian Xu, Ziqiang Feng, Eric Lo
Pages: 1263-1278
doi>10.1145/2882903.2915205
Full text: PDFPDF

Sorting is a crucial operation that could be used to implement SQL operators such as GROUP BY, ORDER BY, and SQL:2003 PARTITION BY. Queries with multiple attributes in those clauses are common in real workloads. When executing queries of that kind, state-of-the-art ...
expand
Elastic Pipelining in an In-Memory Database Cluster
Li Wang, Minqi Zhou, Zhenjie Zhang, Yin Yang, Aoying Zhou, Dina Bitton
Pages: 1279-1294
doi>10.1145/2882903.2882904
Full text: PDFPDF

An in-memory database cluster consists of multiple interconnected nodes with a large capacity of RAM and modern multi-core CPUs. As a conventional query processing strategy, pipelining remains a promising solution for in-memory parallel database systems, ...
expand
Page As You Go: Piecewise Columnar Access In SAP HANA
Reza Sherkat, Colin Florendo, Mihnea Andrei, Anil K. Goel, Anisoara Nica, Peter Bumbulis, Ivan Schreter, Günter Radestock, Christian Bensberg, Daniel Booss, Heiko Gerwens
Pages: 1295-1306
doi>10.1145/2882903.2903729
Full text: PDFPDF

In-memory columnar databases such as SAP HANA achieve extreme performance by means of vector processing over logical units of main memory resident columns. The core in-memory algorithms can be challenged when the working set of an application does not ...
expand
Hybrid Garbage Collection for Multi-Version Concurrency Control in SAP HANA
Juchang Lee, Hyungyu Shin, Chang Gyoo Park, Seongyun Ko, Jaeyun Noh, Yongjae Chuh, Wolfgang Stephan, Wook-Shin Han
Pages: 1307-1318
doi>10.1145/2882903.2903734
Full text: PDFPDF

While multi-version concurrency control (MVCC) supports fast and robust performance in in-memory, relational databases, it has the potential problem of a growing number of versions over time due to obsolete versions. Although a few TB of main memory ...
expand
UpBit: Scalable In-Memory Updatable Bitmap Indexing
Manos Athanassoulis, Zheng Yan, Stratos Idreos
Pages: 1319-1332
doi>10.1145/2882903.2915964
Full text: PDFPDF

Bitmap indexes are widely used in both scientific and commercial databases. They bring fast read performance for specific types of queries, such as equality and selective range queries. A major drawback of bitmap indexes, however, is that supporting ...
expand
SESSION: Session 15 - Interactive Analytics
FluxQuery: An Execution Framework for Highly Interactive Query Workloads
Roee Ebenstein, Niranjan Kamat, Arnab Nandi
Pages: 1333-1345
doi>10.1145/2882903.2882945
Full text: PDFPDF

Modern computing devices and user interfaces have necessitated highly interactive querying. Some of these interfaces issue a large number of dynamically changing and continuous queries to the backend. In others, users expect to inspect results during ...
expand
iOLAP: Managing Uncertainty for Efficient Incremental OLAP
Kai Zeng, Sameer Agarwal, Ion Stoica
Pages: 1347-1361
doi>10.1145/2882903.2915240
Full text: PDFPDF

The size of data and the complexity of analytics continue to grow along with the need for timely and cost-effective analysis. However, the growth of computation power cannot keep up with the growth of data. This calls for a paradigm shift from traditional ...
expand
Dynamic Prefetching of Data Tiles for Interactive Visualization
Leilani Battle, Remco Chang, Michael Stonebraker
Pages: 1363-1375
doi>10.1145/2882903.2882919
Full text: PDFPDF

In this paper, we present ForeCache, a general-purpose tool for exploratory browsing of large datasets. ForeCache utilizes a client-server architecture, where the user interacts with a lightweight client-side interface to browse datasets, and the data ...
expand
Expressive Query Construction through Direct Manipulation of Nested Relational Results
Eirik Bakke, David R. Karger
Pages: 1377-1392
doi>10.1145/2882903.2915210
Full text: PDFPDF

Despite extensive research on visual query systems, the standard way to interact with relational databases remains to be through SQL queries and tailored form interfaces. We consider three requirements to be essential to a successful alternative: (1) ...
expand
Shasta: Interactive Reporting At Scale
Gokul Nath Babu Manoharan, Stephan Ellner, Karl Schnaitter, Sridatta Chegu, Alejandro Estrella-Balderrama, Stephan Gudmundson, Apurv Gupta, Ben Handy, Bart Samwel, Chad Whipkey, Larysa Aharkava, Himani Apte, Nitin Gangahar, Jun Xu, Shivakumar Venkataraman, Divyakant Agrawal, Jeffrey D. Ullman
Pages: 1393-1404
doi>10.1145/2882903.2904444
Full text: PDFPDF

We describe Shasta, a middleware system built at Google to support interactive reporting in complex user-facing applications related to Google's Internet advertising business. Shasta targets applications with challenging requirements: First, user query ...
expand
Datometry Hyper-Q: Bridging the Gap Between Real-Time and Historical Analytics
Lyublena Antova, Rhonda Baldwin, Derrick Bryant, Tuan Cao, Michael Duller, John Eshleman, Zhongxian Gu, Entong Shen, Mohamed A. Soliman, F. Michael Waas
Pages: 1405-1416
doi>10.1145/2882903.2903739
Full text: PDFPDF

Wall Street's trading engines are complex database applications written for time series databases like kdb+ that uses the query language Q to perform real-time analysis. Extending the models to include other data sources, e.g., historic data, is critical ...
expand
SESSION: Session 16 - Streaming 2: Sketches
Time Adaptive Sketches (Ada-Sketches) for Summarizing Data Streams
Anshumali Shrivastava, Arnd Christian Konig, Mikhail Bilenko
Pages: 1417-1432
doi>10.1145/2882903.2882946
Full text: PDFPDF

Obtaining frequency information of data streams, in limited space, is a well-recognized problem in literature. A number of recent practical applications (such as those in computational advertising) require temporally-aware solutions: obtaining historical ...
expand
Streaming Algorithms for Robust Distinct Elements
Di Chen, Qin Zhang
Pages: 1433-1447
doi>10.1145/2882903.2882915
Full text: PDFPDF

We study the problem of estimating distinct elements in the data stream model, which has a central role in traffic monitoring, query optimization, data mining and data integration. Different from all previous work, we study the problem in the noisy ...
expand
Augmented Sketch: Faster and More Accurate Stream Processing
Pratanu Roy, Arijit Khan, Gustavo Alonso
Pages: 1449-1463
doi>10.1145/2882903.2882948
Full text: PDFPDF

Approximated algorithms are often used to estimate the frequency of items on high volume, fast data streams. The most common ones are variations of Count-Min sketch, which use sub-linear space for the count, but can produce errors in the counts of the ...
expand
Matrix Sketching Over Sliding Windows
Zhewei Wei, Xuancheng Liu, Feifei Li, Shuo Shang, Xiaoyong Du, Ji-Rong Wen
Pages: 1465-1480
doi>10.1145/2882903.2915228
Full text: PDFPDF

Large-scale matrix computation becomes essential for many data data applications, and hence the problem of sketching matrix with small space and high precision has received extensive study for the past few years. This problem is often considered in the ...
expand
Graph Stream Summarization: From Big Bang to Big Crunch
Nan Tang, Qing Chen, Prasenjit Mitra
Pages: 1481-1496
doi>10.1145/2882903.2915223
Full text: PDFPDF

A graph stream, which refers to the graph with edges being updated sequentially in a form of a stream, has important applications in cyber security and social networks. Due to the sheer volume and highly dynamic nature of graph streams, the practical ...
expand
Scalable Approximate Query Tracking over Highly Distributed Data Streams
Nikos Giatrakos, Antonios Deligiannakis, Minos Garofalakis
Pages: 1497-1512
doi>10.1145/2882903.2915225
Full text: PDFPDF

The recently-proposed Geometric Monitoring (GM) method has provided a general tool for the distributed monitoring of arbitrary non-linear queries over streaming data observed by a collection of remote sites, with numerous practical applications. Unfortunately, ...
expand
SESSION: Session 17 - Transaction Processing
A Hybrid B+-tree as Solution for In-Memory Indexing on CPU-GPU Heterogeneous Computing Platforms
Amirhesam Shahvarani, Hans-Arno Jacobsen
Pages: 1523-1538
doi>10.1145/2882903.2882918
Full text: PDFPDF

An in-memory indexing tree is a critical component of many databases. Modern many-core processors, such as GPUs, are offering tremendous amounts of computing power making them an attractive choice for accelerating indexing. However, the memory available ...
expand
Low-Overhead Asynchronous Checkpointing in Main-Memory Database Systems
Kun Ren, Thaddeus Diamond, Daniel J. Abadi, Alexander Thomson
Pages: 1539-1551
doi>10.1145/2882903.2915966
Full text: PDFPDF

As it becomes increasingly common for transaction processing systems to operate on datasets that fit within the main memory of a single machine or a cluster of commodity machines, traditional mechanisms for guaranteeing transaction durability---which ...
expand
T-Part: Partitioning of Transactions for Forward-Pushing in Deterministic Database Systems
Shan-Hung Wu, Tsai-Yu Feng, Meng-Kai Liao, Shao-Kan Pi, Yu-Shan Lin
Pages: 1553-1565
doi>10.1145/2882903.2915227
Full text: PDFPDF

Deterministic database systems have been shown to yield high throughput on a cluster of commodity machines while ensuring the strong consistency between replicas, provided that the data can be well-partitioned on these machines. However, data partitioning ...
expand
Reducing the Storage Overhead of Main-Memory OLTP Databases with Hybrid Indexes
Huanchen Zhang, David G. Andersen, Andrew Pavlo, Michael Kaminsky, Lin Ma, Rui Shen
Pages: 1567-1581
doi>10.1145/2882903.2915222
Full text: PDFPDF

Using indexes for query execution is crucial for achieving high performance in modern on-line transaction processing databases. For a main-memory database, however, these indexes consume a large fraction of the total memory available and are thus a major ...
expand
Design Principles for Scaling Multi-core OLTP Under High Contention
Kun Ren, Jose M. Faleiro, Daniel J. Abadi
Pages: 1583-1598
doi>10.1145/2882903.2882958
Full text: PDFPDF

Although significant recent progress has been made in improving the multi-core scalability of high throughput transactional database systems, modern systems still fail to achieve scalable throughput for workloads involving frequent access to highly contended ...
expand
DBSherlock: A Performance Diagnostic Tool for Transactional Databases
Dong Young Yoon, Ning Niu, Barzan Mozafari
Pages: 1599-1614
doi>10.1145/2882903.2915218
Full text: PDFPDF

Running an online transaction processing (OLTP) system is one of the most daunting tasks required of database administrators (DBAs). As businesses rely on OLTP databases to support their mission-critical and real-time applications, poor database performance ...
expand
SESSION: Session 18 - Transactions and Consistency
TARDiS: A Branch-and-Merge Approach To Weak Consistency
Natacha Crooks, Youer Pu, Nancy Estrada, Trinabh Gupta, Lorenzo Alvisi, Allen Clement
Pages: 1615-1628
doi>10.1145/2882903.2882951
Full text: PDFPDF

This paper presents the design, implementation, and evaluation of TARDiS (Transactional Asynchronously Replicated Divergent Store), a transactional key-value store explicitly designed for weakly-consistent systems. Reasoning about these systems is hard, ...
expand
TicToc: Time Traveling Optimistic Concurrency Control
Xiangyao Yu, Andrew Pavlo, Daniel Sanchez, Srinivas Devadas
Pages: 1629-1642
doi>10.1145/2882903.2882935
Full text: PDFPDF

Concurrency control for on-line transaction processing (OLTP) database management systems (DBMSs) is a nasty game. Achieving higher performance on emerging many-core systems is difficult. Previous research has shown that timestamp management is the key ...
expand
Scaling Multicore Databases via Constrained Parallel Execution
Zhaoguo Wang, Shuai Mu, Yang Cui, Han Yi, Haibo Chen, Jinyang Li
Pages: 1643-1658
doi>10.1145/2882903.2882934
Full text: PDFPDF

Multicore in-memory databases often rely on traditional con- currency control schemes such as two-phase-locking (2PL) or optimistic concurrency control (OCC). Unfortunately, when the workload exhibits a non-trivial amount of contention, both 2PL and ...
expand
Towards a Non-2PC Transaction Management in Distributed Database Systems
Qian Lin, Pengfei Chang, Gang Chen, Beng Chin Ooi, Kian-Lee Tan, Zhengkui Wang
Pages: 1659-1674
doi>10.1145/2882903.2882923
Full text: PDFPDF

Shared-nothing architecture has been widely used in distributed databases to achieve good scalability. While it offers superior performance for local transactions, the overhead of processing distributed transactions can degrade the system performance ...
expand
ERMIA: Fast Memory-Optimized Database System for Heterogeneous Workloads
Kangnyeon Kim, Tianzheng Wang, Ryan Johnson, Ippokratis Pandis
Pages: 1675-1687
doi>10.1145/2882903.2882905
Full text: PDFPDF

Large main memories and massively parallel processors have triggered not only a resurgence of high-performance transaction processing systems optimized for large main-memory and massively parallel processors, but also an increasing demand for processing ...
expand
Transaction Healing: Scaling Optimistic Concurrency Control on Multicores
Yingjun Wu, Chee-Yong Chan, Kian-Lee Tan
Pages: 1689-1704
doi>10.1145/2882903.2915202
Full text: PDFPDF

Today's main-memory databases can support very high transaction rate for OLTP applications. However, when a large number of concurrent transactions contend on the same data records, the system performance can deteriorate significantly. This is especially ...
expand
SESSION: Session 19 - Query Optimization
Enabling Incremental Query Re-Optimization
Mengmeng Liu, Zachary G. Ives, Boon Thau Loo
Pages: 1705-1720
doi>10.1145/2882903.2915212
Full text: PDFPDF

As declarative query processing techniques expand to the Web, data streams, network routers, and cloud platforms, there is an increasing need to re-plan execution in the presence of unanticipated performance changes. New runtime information may affect ...
expand
Sampling-Based Query Re-Optimization
Wentao Wu, Jeffrey F. Naughton, Harneet Singh
Pages: 1721-1736
doi>10.1145/2882903.2882914
Full text: PDFPDF

Despite of decades of work, query optimizers still make mistakes on "difficult" queries because of bad cardinality estimates, often due to the interaction of multiple predicates and correlations in the data. In this paper, we propose a low-cost post-processing ...
expand
A Fast Randomized Algorithm for Multi-Objective Query Optimization
Immanuel Trummer, Christoph Koch
Pages: 1737-1752
doi>10.1145/2882903.2882927
Full text: PDFPDF

Query plans are compared according to multiple cost metrics in multi-objective query optimization. The goal is to find the set of Pareto plans realizing optimal cost tradeoffs for a given query. So far, only algorithms with exponential complexity in ...
expand
Operator and Query Progress Estimation in Microsoft SQL Server Live Query Statistics
Kukjin Lee, Arnd Christian König, Vivek Narasayya, Bolin Ding, Surajit Chaudhuri, Brent Ellwein, Alexey Eksarevskiy, Manbeen Kohli, Jacob Wyant, Praneeta Prakash, Rimma Nehme, Jiexing Li, Jeff Naughton
Pages: 1753-1764
doi>10.1145/2882903.2903728
Full text: PDFPDF

We describe the design and implementation of the new Live Query Statistics (LQS) feature in Microsoft SQL Server 2016. The functionality includes the display of overall query progress as well as progress of individual operators in the query execution ...
expand
Optimization of Nested Queries using the NF2 Algebra
Jürgen Hölsch, Michael Grossniklaus, Marc H. Scholl
Pages: 1765-1780
doi>10.1145/2882903.2915241
Full text: PDFPDF

A key promise of SQL is that the optimizer will find the most efficient execution plan, regardless of how the query is formulated. In general, query optimizers of modern database systems are able to keep this promise, with the notable exception of nested ...
expand
Extracting Equivalent SQL from Imperative Code in Database Applications
K. Venkatesh Emani, Karthik Ramachandra, Subhro Bhattacharya, S. Sudarshan
Pages: 1781-1796
doi>10.1145/2882903.2882926
Full text: PDFPDF

Optimizing the performance of database applications is an area of practical importance, and has received significant attention in recent years. In this paper we present an approach to this problem which is based on extracting a concise algebraic representation ...
expand
SESSION: Session 20 - Graphs 3: Potpourri
Generating Preview Tables for Entity Graphs
Ning Yan, Sona Hasani, Abolfazl Asudeh, Chengkai Li
Pages: 1797-1811
doi>10.1145/2882903.2915221
Full text: PDFPDF

Users are tapping into massive, heterogeneous entity graphs for many applications. It is challenging to select entity graphs for a particular need, given abundant datasets from many sources and the oftentimes scarce information for them. We propose methods ...
expand
Speedup Graph Processing by Graph Ordering
Hao Wei, Jeffrey Xu Yu, Can Lu, Xuemin Lin
Pages: 1813-1828
doi>10.1145/2882903.2915220
Full text: PDFPDF

The CPU cache performance is one of the key issues to efficiency in database systems. It is reported that cache miss latency takes a half of the execution time in database systems. To improve the CPU cache performance, there are studies to support searching ...
expand
ROLL: Fast In-Memory Generation of Gigantic Scale-free Networks
Ali Hadian, Sadegh Nobari, Behrooz Minaei-Bidgoli, Qiang Qu
Pages: 1829-1842
doi>10.1145/2882903.2882964
Full text: PDFPDF

Real-world graphs are not always publicly available or sometimes do not meet specific research requirements. These challenges call for generating synthetic networks that follow properties of the real-world networks. Barabási-Albert (BA) is a well-known ...
expand
Functional Dependencies for Graphs
Wenfei Fan, Yinghui Wu, Jingbo Xu
Pages: 1843-1857
doi>10.1145/2882903.2915232
Full text: PDFPDF

We propose a class of functional dependencies for graphs, referred to as GFDs. GFDs capture both attribute-value dependencies and topological structures of entities, and subsume conditional functional dependencies (CFDs) as a special case. We show that ...
expand
SLING: A Near-Optimal Index Structure for SimRank
Boyu Tian, Xiaokui Xiao
Pages: 1859-1874
doi>10.1145/2882903.2915243
Full text: PDFPDF

SimRank is a similarity measure for graph nodes that has numerous applications in practice. Scalable SimRank computation has been the subject of extensive research for more than a decade, and yet, none of the existing solutions can efficiently derive ...
expand
Query Planning for Evaluating SPARQL Property Paths
Nikolay Yakovets, Parke Godfrey, Jarek Gryz
Pages: 1875-1889
doi>10.1145/2882903.2882944
Full text: PDFPDF

The extension of SPARQL in version 1.1 with property paths offers a type of regular path query for RDF graph databases. Such queries are difficult to optimize and evaluate efficiently, however. We have embarked on a project, Waveguide, ...
expand
SESSION: Session 21 - Hardware Acceleration and Query Compilation
Robust Query Processing in Co-Processor-accelerated Databases
Sebastian Breß, Henning Funke, Jens Teubner
Pages: 1891-1906
doi>10.1145/2882903.2882936
Full text: PDFPDF

Technology limitations are making the use of heterogeneous computing devices much more than an academic curiosity. In fact, the use of such devices is widely acknowledged to be the only promising way to achieve application-speedups that users urgently ...
expand
How to Architect a Query Compiler
Amir Shaikhha, Yannis Klonatos, Lionel Parreaux, Lewis Brown, Mohammad Dashti, Christoph Koch
Pages: 1907-1922
doi>10.1145/2882903.2915244
Full text: PDFPDF

This paper studies architecting query compilers. The state of the art in query compiler construction is lagging behind that in the compilers field. We attempt to remedy this by exploring the key causes of technical challenges in need of well founded ...
expand
Automated Demand-driven Resource Scaling in Relational Database-as-a-Service
Sudipto Das, Feng Li, Vivek R. Narasayya, Arnd Christian König
Pages: 1923-1934
doi>10.1145/2882903.2903733
Full text: PDFPDF

Relational Database-as-a-Service (DaaS) platforms today support the abstraction of a resource container that guarantees a fixed amount of resources. Tenants are responsible for selecting a container size suitable for their workloads, which they can change ...
expand
GPL: A GPU-based Pipelined Query Processing Engine
Johns Paul, Jiong He, Bingsheng He
Pages: 1935-1950
doi>10.1145/2882903.2915224
Full text: PDFPDF

Graphics Processing Units (GPUs) have evolved as a powerful query co-processor for main memory On-Line Analytical Processing (OLAP) databases. However, existing GPU-based query processors adopt a kernel-based execution approach which optimizes individual ...
expand
Towards a Hybrid Design for Fast Query Processing in DB2 with BLU Acceleration Using Graphical Processing Units: A Technology Demonstration
Sina Meraji, Berni Schiefer, Lan Pham, Lee Chu, Peter Kokosielis, Adam Storm, Wayne Young, Chang Ge, Geoffrey Ng, Kajan Kanagaratnam
Pages: 1951-1960
doi>10.1145/2882903.2903735
Full text: PDFPDF

In this paper, we show how we use Nvidia GPUs and host CPU cores for faster query processing in a DB2 database using BLU Acceleration (DB2's column store technology). Moreover, we show the benefits and problems of using hardware accelerators (more specifically ...
expand
An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory
Stefan Schuh, Xiao Chen, Jens Dittrich
Pages: 1961-1976
doi>10.1145/2882903.2882917
Full text: PDFPDF

Relational equi-joins are at the heart of almost every query plan. They have been studied, improved, and reexamined on a regular basis since the existence of the database community. In the past four years several new join algorithms have been proposed ...
expand
SESSION: Session 22 - Nearest Neighbors and Similarity Search
Top-k Relevant Semantic Place Retrieval on Spatial RDF Data
Jieming Shi, Dingming Wu, Nikos Mamoulis
Pages: 1977-1990
doi>10.1145/2882903.2882941
Full text: PDFPDF

RDF data are traditionally accessed using structured query languages, such as SPARQL. However, this requires users to understand the language as well as the RDF schema. Keyword search on RDF data aims at relieving the user from these requirements; the ...
expand
Local Similarity Search for Unstructured Text
Pei Wang, Chuan Xiao, Jianbin Qin, Wei Wang, Xiaoyang Zhang, Yoshiharu Ishikawa
Pages: 1991-2005
doi>10.1145/2882903.2915211
Full text: PDFPDF

With the growing popularity of electronic documents, replication can occur for many reasons. People may copy text segments from various sources and make modifications. In this paper, we study the problem of local similarity search to find partially replicated ...
expand
Similarity Join over Array Data
Weijie Zhao, Florin Rusu, Bin Dong, Kesheng Wu
Pages: 2007-2022
doi>10.1145/2882903.2915247
Full text: PDFPDF

Scientific applications are generating an ever-increasing volume of multi-dimensional data that are largely processed inside distributed array databases and frameworks. Similarity join is a fundamental operation across scientific workloads that requires ...
expand
LazyLSH: Approximate Nearest Neighbor Search for Multiple Distance Functions with a Single Index
Yuxin Zheng, Qi Guo, Anthony K.H. Tung, Sai Wu
Pages: 2023-2037
doi>10.1145/2882903.2882930
Full text: PDFPDF

Due to the "curse of dimensionality" problem, it is very expensive to process the nearest neighbor (NN) query in high-dimensional spaces; and hence, approximate approaches, such as Locality-Sensitive Hashing (LSH), are widely used for their theoretical ...
expand
Set-based Similarity Search for Time Series
Jinglin Peng, Hongzhi Wang, Jianzhong Li, Hong Gao
Pages: 2039-2052
doi>10.1145/2882903.2882963
Full text: PDFPDF

A fundamental problem of time series is k nearest neighbor (k-NN) query processing. However, existing methods are not fast enough for large dataset. In this paper, we propose a novel approach, STS3, to process k-NN queries by transforming time series ...
expand
Range-based Obstructed Nearest Neighbor Queries
Huaijie Zhu, Xiaochun Yang, Bin Wang, Wang-Chien Lee
Pages: 2053-2068
doi>10.1145/2882903.2915234
Full text: PDFPDF

In this paper, we study a novel variant of obstructed nearest neighbor queries, namely, range-based obstructed nearest neighbor (RONN) search. A natural generalization of continuous obstructed nearest-neighbor (CONN), an RONN query retrieves ...
expand
DEMONSTRATION SESSION: Session 23 - Demonstrations
Rheem: Enabling Multi-Platform Task Execution
Divy Agrawal, Lamine Ba, Laure Berti-Equille, Sanjay Chawla, Ahmed Elmagarmid, Hossam Hammady, Yasser Idris, Zoi Kaoudi, Zuhair Khayyat, Sebastian Kruse, Mourad Ouzzani, Paolo Papotti, Jorge-Arnulfo Quiane-Ruiz, Nan Tang, Mohammed J. Zaki
Pages: 2069-2072
doi>10.1145/2882903.2899414
Full text: PDFPDF

Many emerging applications, from domains such as healthcare and oil & gas, require several data processing systems for complex analytics. This demo paper showcases system, a framework that provides multi-platform task execution for such applications. ...
expand
Emma in Action: Declarative Dataflows for Scalable Data Analysis
Alexander Alexandrov, Andreas Salzmann, Georgi Krastev, Asterios Katsifodimos, Volker Markl
Pages: 2073-2076
doi>10.1145/2882903.2899396
Full text: PDFPDF

Parallel dataflow APIs based on second-order functions were originally seen as a flexible alternative to SQL. Over time, however, their complexity increased due to the number of physical aspects that had to be exposed by the underlying engines in order ...
expand
Wildfire: Concurrent Blazing Data Ingest and Analytics
Ronald Barber, Matt Huras, Guy Lohman, C. Mohan, Rene Mueller, Fatma Özcan, Hamid Pirahesh, Vijayshankar Raman, Richard Sidle, Oleg Sidorkin, Adam Storm, Yuanyuan Tian, Pinar Tözun
Pages: 2077-2080
doi>10.1145/2882903.2899406
Full text: PDFPDF

We demonstrate Hybrid Transactional and Analytics Processing (HTAP) on the Spark platform by the Wildfire prototype, which can ingest up to ~6 million inserts per second per node and simultaneously perform complex SQL analytics queries. Here, a simplified ...
expand
Efficient Query Processing on Many-core Architectures: A Case Study with Intel Xeon Phi Processor
Xuntao Cheng, Bingsheng He, Mian Lu, Chiew Tong Lau, Huynh Phung Huynh, Rick Siow Mong Goh
Pages: 2081-2084
doi>10.1145/2882903.2899407
Full text: PDFPDF

Recently, Intel Xeon Phi is emerging as a many-core processor with up to 61 x86 cores. In this demonstration, we present PhiDB, an OLAP query processor with simultaneous multi-threading (SMT) capabilities on Xeon Phi as a case study for parallel database ...
expand
ReproZip: Computational Reproducibility With Ease
Fernando Chirigati, Rémi Rampin, Dennis Shasha, Juliana Freire
Pages: 2085-2088
doi>10.1145/2882903.2899401
Full text: PDFPDF

We present ReproZip, the recommended packaging tool for the SIGMOD Reproducibility Review. ReproZip was designed to simplify the process of making an existing computational experiment reproducible across platforms, even when the experiment was put together ...
expand
CLAMS: Bringing Quality to Data Lakes
Mina Farid, Alexandra Roatis, Ihab F. Ilyas, Hella-Franziska Hoffmann, Xu Chu
Pages: 2089-2092
doi>10.1145/2882903.2899391
Full text: PDFPDF

With the increasing incentive of enterprises to ingest as much data as they can in what is commonly referred to as "data lakes", and with the recent development of multiple technologies to support this "load-first" paradigm, the new environment presents ...
expand
FERARI: A Prototype for Complex Event Processing over Streaming Multi-cloud Platforms
Ioannis Flouris, Vasiliki Manikaki, Nikos Giatrakos, Antonios Deligiannakis, Minos Garofalakis, Michael Mock, Sebastian Bothe, Inna Skarbovsky, Fabiana Fournier, Marko Stajcer, Tomislav Krizan, Jonathan Yom-Tov, Taji Curin
Pages: 2093-2096
doi>10.1145/2882903.2899395
Full text: PDFPDF

In this demo, we present FERARI, a prototype that enables real-time Complex Event Processing (CEP) for large volume event data streams over distributed topologies. Our prototype constitutes, to our knowledge, the first complete, multi-cloud based end-to-end ...
expand
Constance: An Intelligent Data Lake System
Rihan Hai, Sandra Geisler, Christoph Quix
Pages: 2097-2100
doi>10.1145/2882903.2899389
Full text: PDFPDF

As the challenge of our time, Big Data still has many research hassles, especially the variety of data. The high diversity of data sources often results in information silos, a collection of non-integrated data management systems with heterogeneous schemas, ...
expand
Exploring Privacy-Accuracy Tradeoffs using DPComp
Michael Hay, Ashwin Machanavajjhala, Gerome Miklau, Yan Chen, Dan Zhang, George Bissias
Pages: 2101-2104
doi>10.1145/2882903.2899387
Full text: PDFPDF

The emergence of differential privacy as a primary standard for privacy protection has led to the development, by the research community, of hundreds of algorithms for various data analysis tasks. Yet deployment of these techniques has been slowed by ...
expand
Interactive Search and Exploration of Waveform Data with Searchlight
Alexander Kalinin, Ugur Cetintemel, Stan Zdonik
Pages: 2105-2108
doi>10.1145/2882903.2899404
Full text: PDFPDF

Searchlight enables search and exploration of large, multi-dimensional data sets interactively. It allows users to explore by specifying rich constraints for the "objects" they are interested in identifying. Constraints can express a variety of properties, ...
expand
Ontology-Based Integration of Streaming and Static Relational Data with Optique
Evgeny Kharlamov, Sebastian Brandt, Ernesto Jimenez-Ruiz, Yannis Kotidis, Steffen Lamparter, Theofilos Mailis, Christian Neuenstadt, Özgür Özçep, Christoph Pinkel, Christoforos Svingos, Dmitriy Zheleznyakov, Ian Horrocks, Yannis Ioannidis, Ralf Moeller
Pages: 2109-2112
doi>10.1145/2882903.2899385
Full text: PDFPDF

Real-time processing of data coming from multiple heterogeneous data streams and static databases is a typical task in many industrial scenarios such as diagnostics of large machines. A complex diagnostic task may require a collection of up to hundreds ...
expand
The CloudMdsQL Multistore System
Boyan Kolev, Carlyna Bondiombouy, Patrick Valduriez, Ricardo Jimenez-Peris, Raquel Pau, José Pereira
Pages: 2113-2116
doi>10.1145/2882903.2899400
Full text: PDFPDF

The blooming of different cloud data management infrastructures has turned multistore systems to a major topic in the nowadays cloud landscape. In this demonstration, we present a Cloud Multidatastore Query Language (CloudMdsQL), and its query engine. ...
expand
ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning
Sanjay Krishnan, Michael J. Franklin, Ken Goldberg, Jiannan Wang, Eugene Wu
Pages: 2117-2120
doi>10.1145/2882903.2899409
Full text: PDFPDF

Databases can be corrupted with various errors such as missing, incorrect, or inconsistent values. Increasingly, modern data analysis pipelines involve Machine Learning, and the effects of dirty data can be difficult to debug.Dirty data is often sparse, ...
expand
Wander Join: Online Aggregation for Joins
Feifei Li, Bin Wu, Ke Yi, Zhuoyue Zhao
Pages: 2121-2124
doi>10.1145/2882903.2899413
Full text: PDFPDF

Joins are expensive, and online aggregation over joins was proposed to mitigate the cost, which offers a nice and flexible tradeoff between query efficiency and accuracy in a continuous, online fashion. However, the state-of-the-art approach, in both ...
expand
PerNav: A Route Summarization Framework for Personalized Navigation
Yaguang Li, Han Su, Ugur Demiryurek, Bolong Zheng, Kai Zeng, Cyrus Shahabi
Pages: 2125-2128
doi>10.1145/2882903.2899384
Full text: PDFPDF

In this paper, we study a route summarization framework for Personalized Navigation dubbed PerNav - with which the goal is to generate more intuitive and customized turn-by-turn directions based on user generated content. The turn-by-turn directions ...
expand
Making the Case for Query-by-Voice with EchoQuery
Gabriel Lyons, Vinh Tran, Carsten Binnig, Ugur Cetintemel, Tim Kraska
Pages: 2129-2132
doi>10.1145/2882903.2899394
Full text: PDFPDF

Recent advances in automatic speech recognition and natural language processing have led to a new generation of robust voice-based interfaces. Yet, there is very little work on using voice-based interfaces to query database systems. In fact, one might ...
expand
QUEPA: QUerying and Exploring a Polystore by Augmentation
Antonio Maccioni, Edoardo Basili, Riccardo Torlone
Pages: 2133-2136
doi>10.1145/2882903.2899393
Full text: PDFPDF

Polystore systems (or simply polystores) have been recently proposed to support a common scenario in which enterprise data are stored in a variety of database technologies relying on different data models and languages. Polystores provide a loosely coupled ...
expand
REACT: Context-Sensitive Recommendations for Data Analysis
Tova Milo, Amit Somech
Pages: 2137-2140
doi>10.1145/2882903.2899392
Full text: PDFPDF

Data analysis may be a difficult task, especially for non-expert users, as it requires deep understanding of the investigated domain and the particular context. In this demo we present REACT, a system that hooks to the analysis UI and provides the users ...
expand
PerfEnforce Demonstration: Data Analytics with Performance Guarantees
Jennifer Ortiz, Brendan Lee, Magdalena Balazinska
Pages: 2141-2144
doi>10.1145/2882903.2899402
Full text: PDFPDF

We demonstrate PerfEnforce, a dynamic scaling engine for analytics services. PerfEnforce automatically scales a cluster of virtual machines in order to minimize costs while probabilistically meeting the query runtime guarantees offered by a performance-oriented ...
expand
High-Performance Geospatial Analytics in HyPerSpace
Varun Pandey, Andreas Kipf, Dimitri Vorona, Tobias Mühlbauer, Thomas Neumann, Alfons Kemper
Pages: 2145-2148
doi>10.1145/2882903.2899412
Full text: PDFPDF

In the past few years, massive amounts of location-based data has been captured. Numerous datasets containing user location information are readily available to the public. Analyzing such datasets can lead to fascinating insights into the mobility patterns ...
expand
What Makes a Good Physical plan?: Experiencing Hardware-Conscious Query Optimization with Candomblé
Holger Pirk, Oscar Moll, Sam Madden
Pages: 2149-2152
doi>10.1145/2882903.2899410
Full text: PDFPDF

Query optimization is hard and the current proliferation of "modern" hardware does nothing to make it any easier. In addition, the tools that are commonly used by performance engineers, such as compiler intrinsics, static analyzers or hardware performance ...
expand
SnappyData: A Hybrid Transactional Analytical Store Built On Spark
Jags Ramnarayan, Barzan Mozafari, Sumedh Wale, Sudhir Menon, Neeraj Kumar, Hemant Bhanawat, Soubhik Chakraborty, Yogesh Mahajan, Rishitesh Mishra, Kishor Bachhav
Pages: 2153-2156
doi>10.1145/2882903.2899408
Full text: PDFPDF

In recent years, our customers have expressed frustration in the traditional approach of using a combination of disparate products to handle their streaming, transactional and analytical needs. The common practice of stitching heterogeneous environments ...
expand
SourceSight: Enabling Effective Source Selection
Theodoros Rekatsinas, Amol Deshpande, Xin Luna Dong, Lise Getoor, Divesh Srivastava
Pages: 2157-2160
doi>10.1145/2882903.2899403
Full text: PDFPDF

Recently there has been a rapid increase in the number of data sources and data services, such as cloud-based data markets and data portals, that facilitate the collection, publishing and trading of data. Data sources typically exhibit large heterogeneity ...
expand
BART in Action: Error Generation and Empirical Evaluations of Data-Cleaning Systems
Donatello Santoro, Patricia C. Arocena, Boris Glavic, Giansalvatore Mecca, Renée J. Miller, Paolo Papotti
Pages: 2161-2164
doi>10.1145/2882903.2899397
Full text: PDFPDF

Repairing erroneous or conflicting data that violate a set of constraints is an important problem in data management. Many automatic or semi-automatic data-repairing algorithms have been proposed in the last few years, each with its own strengths and ...
expand
RxSpatial: Reactive Spatial Library for Real-Time Location Tracking and Processing
Youying Shi, Abdeltawab M. Hendawi, Hossam Fattah, Mohamed Ali
Pages: 2165-2168
doi>10.1145/2882903.2899411
Full text: PDFPDF

Current commercial spatial libraries implemented strong support on functionalities like intersection, distance, and area for various stationary geospatial objects. The missing point is the support for moving object. Performing moving object real-time ...
expand
Web-based Benchmarks for Forecasting Systems: The ECAST Platform
Robert Ulbricht, Claudio Hartmann, Martin Hahmann, Hilko Donker, Wolfgang Lehner
Pages: 2169-2172
doi>10.1145/2882903.2899399
Full text: PDFPDF

The role of precise forecasts in the energy domain has changed dramatically. New supply forecasting methods are developed to better address this challenge, but meaningful benchmarks are rare and time-intensive. We propose the ECAST online platform in ...
expand
Energy Elasticity on Heterogeneous Hardware using Adaptive Resource Reconfiguration LIVE
Annett Ungethüm, Thomas Kissinger, Willi-Wolfram Mentzel, Dirk Habich, Wolfgang Lehner
Pages: 2173-2176
doi>10.1145/2882903.2899390
Full text: PDFPDF

Energy awareness of database systems has emerged as a critical research topic, since energy consumption is becoming a major limiter for their scalability. Recent energy-related hardware developments trend towards offering more and more configuration ...
expand
QFix: Demonstrating Error Diagnosis in Query Histories
Xiaolan Wang, Alexandra Meliou, Eugene Wu
Pages: 2177-2180
doi>10.1145/2882903.2899388
Full text: PDFPDF

An increasing number of applications in all aspects of society rely on data. Despite the long line of research in data cleaning and repairs, data correctness has been an elusive goal. Errors in the data can be extremely disruptive, and are detrimental ...
expand
CoDAR: Revealing the Generalized Procedure & Recommending Algorithms of Community Detection
Xiang Ying, Chaokun Wang, Meng Wang, Jeffrey Xu Yu, Jun Zhang
Pages: 2181-2184
doi>10.1145/2882903.2899386
Full text: PDFPDF

Community detection has attracted great interest in graph analysis and mining during the past decade, and a great number of approaches have been developed to address this problem. However, the lack of a uniform framework and a reasonable evaluation method ...
expand
DB-Risk: The Game of Global Database Placement
Victor Zakhary, Faisal Nawab, Divyakant Agrawal, Amr El Abbadi
Pages: 2185-2188
doi>10.1145/2882903.2899405
Full text: PDFPDF

Geo-replication is the process of maintaining copies of data at geographically dispersed datacenters for better availability and fault-tolerance. The distinguishing characteristic of geo-replication is the large wide-area latency between datacenters ...
expand
Quegel: A General-Purpose System for Querying Big Graphs
Qizhen Zhang, Da Yan, James Cheng
Pages: 2189-2192
doi>10.1145/2882903.2899398
Full text: PDFPDF

Inspired by Google's Pregel, many distributed graph processing systems have been developed recently to process big graphs. These systems expose a vertex-centric programming interface to users, where a programmer thinks like a vertex when designing parallel ...
expand
TUTORIAL SESSION: Session 24 - Tutorials
Introduction to Spark 2.0 for Database Researchers
Michael Armbrust, Doug Bateman, Reynold Xin, Matei Zaharia
Pages: 2193-2194
doi>10.1145/2882903.2912565
Full text: PDFPDF

Originally started as an academic research project at UC Berkeley, Apache Spark is one of the most popular open source projects for big data analytics. Over 1000 volunteers have contributed code to the project; it is supported by virtually every commercial ...
expand
Design Tradeoffs of Data Access Methods
Manos Athanassoulis, Stratos Idreos
Pages: 2195-2200
doi>10.1145/2882903.2912569
Full text: PDFPDF

Database researchers and practitioners have been building methods to store, access, and update data for more than five decades. Designing access methods has been a constant effort to adapt to the ever changing underlying hardware and workload requirements. ...
expand
Data Cleaning: Overview and Emerging Challenges
Xu Chu, Ihab F. Ilyas, Sanjay Krishnan, Jiannan Wang
Pages: 2201-2206
doi>10.1145/2882903.2912574
Full text: PDFPDF

Detecting and repairing dirty data is one of the perennial challenges in data analytics, and failure to do so can result in inaccurate analytics and unreliable decisions. Over the past few years, there has been a surge of interest from both industry ...
expand
Querying Geo-Textual Data: Spatial Keyword Queries and Beyond
Gao Cong, Christian S. Jensen
Pages: 2207-2212
doi>10.1145/2882903.2912572
Full text: PDFPDF

Over the past decade, we have moved from a predominantly desktop based web to a predominantly mobile web, where users most often access the web from mobile devices such as smartphones. In addition, we are witnessing a proliferation of geo-located, textual ...
expand
Provenance: On and Behind the Screens
Melanie Herschel, Marcel Hlawatsch
Pages: 2213-2217
doi>10.1145/2882903.2912568
Full text: PDFPDF

Collecting and processing provenance, i.e., information describing the production process of some end product, is important in various applications, e.g., to assess quality, to ensure reproducibility, or to reinforce trust in the end product. In the ...
expand
Microblogs Data Management Systems: Querying, Analysis, and Visualization
Mohamed F. Mokbel, Amr Magdy
Pages: 2219-2222
doi>10.1145/2882903.2912570
Full text: PDFPDF

Microblogs data, e.g., tweets, reviews, news comments, and social media comments, has gained considerable attention in recent years due to its popularity and rich contents. Nowadays, microblogs applications span a wide spectrum of interests, including ...
expand
The Challenges of Global-scale Data Management
Faisal Nawab, Divyakant Agrawal, Amr El Abbadi
Pages: 2223-2227
doi>10.1145/2882903.2912571
Full text: PDFPDF

Global-scale data management (GSDM) empowers systems by providing higher levels of fault-tolerance, read availability, and efficiency in utilizing cloud resources. This has led to the emergence of global-scale data management and event processing. However, ...
expand
Semistructured Models, Queries and Algebras in the Big Data Era: Tutorial Summary
Yannis Papakonstantinou
Pages: 2229-2233
doi>10.1145/2882903.2912573
Full text: PDFPDF

Numerous databases promoted as SQL-on-Hadoop, NewSQL and NoSQL support semi-structured, schemaless and heterogeneous data, typically in the form of enriched JSON. They also provide corresponding query languages. In addition to these genuine JSON databases, ...
expand
Automatic Entity Recognition and Typing in Massive Text Data
Xiang Ren, Ahmed El-Kishky, Heng Ji, Jiawei Han
Pages: 2235-2239
doi>10.1145/2882903.2912567
Full text: PDFPDF

In today's computerized and information-based society, individuals are constantly presented with vast amounts of text data, ranging from news articles, scientific publications, product reviews, to a wide range of textual information from social media. ...
expand
Big Graph Analytics Systems
Da Yan, Yingyi Bu, Yuanyuan Tian, Amol Deshpande, James Cheng
Pages: 2241-2243
doi>10.1145/2882903.2912566
Full text: PDFPDF

In recent years we have witnessed a surging interest in developing Big Graph processing systems. To date, tens of Big Graph systems have been proposed. This tutorial provides a timely and comprehensive review of existing Big Graph systems, and summarizes ...
expand
SESSION: Session 25 - Undergraduate Student Abstracts
Constructing Join Histograms from Histograms with q-error Guarantees
Kaleb Alway, Anisoara Nica
Pages: 2245-2246
doi>10.1145/2882903.2914828
Full text: PDFPDF

Histograms are implemented and used in any database system, usually defined on a single-column of a database table. However, one of the most desired statistical data in such systems are statistics on the correlation among columns. In this paper we present ...
expand
Graph Summarization for Geo-correlated Trends Detection in Social Networks
Colin Biafore, Faisal Nawab
Pages: 2247-2248
doi>10.1145/2882903.2914832
Full text: PDFPDF

Trends detection in social networks is possible via a multitude of models with different characteristics. These models are pre-defined and rigid which creates the need to expose the social network graph to data scientists to introduce the human-element ...
expand
M3: Scaling Up Machine Learning via Memory Mapping
Dezhi Fang, Duen Horng Chau
Pages: 2249-2250
doi>10.1145/2882903.2914830
Full text: PDFPDF

To process data that do not fit in RAM, conventional wisdom would suggest using distributed approaches. However, recent research has demonstrated virtual memory's strong potential in scaling up graph mining algorithms on a single machine. We propose ...
expand
K-means Split Revisited: Well-grounded Approach and Experimental Evaluation
Valentin Grigorev, George Chernishev
Pages: 2251-2252
doi>10.1145/2882903.2914833
Full text: PDFPDF

R-tree is a data structure used for multidimensional indexing. Essentially, it is a balanced tree consisting of nested hyper-rectangles which are used to locate the data. One of the most performance sensitive parts of this data structure is its split ...
expand
Main Memory Adaptive Denormalization
Zezhou Liu, Stratos Idreos
Pages: 2253-2254
doi>10.1145/2882903.2914835
Full text: PDFPDF

Joins have traditionally been the most expensive database operator, but they are required to query normalized schemas. In turn, normalized schemas are necessary to minimize update costs and space usage. Joins can be avoided altogether by using a denormalized ...
expand
Adaptive Data Skipping in Main-Memory Systems
Wilson Qin, Stratos Idreos
Pages: 2255-2256
doi>10.1145/2882903.2914836
Full text: PDFPDF

As modern main-memory optimized data systems increasingly rely on fast scans, lightweight indexes that allow for data skipping play a crucial role in data filtering to reduce system I/O. Scans benefit from data skipping when the data order is sorted, ...
expand
Searching Web Data using MinHash LSH
BiChen Rao, Erkang Zhu
Pages: 2257-2258
doi>10.1145/2882903.2914838
Full text: PDFPDF

In this extended abstract, we explore the use of MinHash Locality Sensitive Hashing (MinHash LSH) to address the problem of indexing and searching Web data. We discuss a statistical tuning strategy of MinHash LSH, and experimentally evaluate the accuracy ...
expand
Research Contribution as a Measure of Influence
Lais M.A. Rocha, Mirella M. Moro
Pages: 2259-2260
doi>10.1145/2882903.2914834
Full text: PDFPDF

We propose the 3c-index that measures the influence degree of researchers by evaluating the links they establish between communities. We evaluate its performance against well known metrics. The results show 3c-index outperforms them in most cases and ...
expand
Vectorizing an In Situ Query Engine
Panagiotis Sioulas, Anastasia Ailamaki
Pages: 2261-2262
doi>10.1145/2882903.2914829
Full text: PDFPDF

Database systems serve a wide range of use cases efficiently, but require data to be loaded and adapted to the system's execution engine. This pre-processing step is a bottleneck to the analysis of the increasingly large and heterogeneous datasets. Therefore, ...
expand
Exploring Visualization of Data Transforms
Larry Xu
Pages: 2263-2264
doi>10.1145/2882903.2914837
Full text: PDFPDF

In the context of data exploration, users often interact with relational database systems in an interactive query session to form useful insights. Each query a user executes can potentially transform a resultset in complex ways. We explore some of the ...
expand
Minimizing Average Regret Ratio in Database
Sepanta Zeighami, Raymond Chi-Wing Wong
Pages: 2265-2266
doi>10.1145/2882903.2914831
Full text: PDFPDF

We propose "average regret ratio" as a metric to measure users' satisfaction after a user sees k selected points of a database, instead of all of the points in the database. We introduce the average regret ratio as another means of multi-criteria decision ...
expand

Powered by The ACM Guide to Computing Literature


The ACM Digital Library is published by the Association for Computing Machinery. Copyright © 2017 ACM, Inc.
Terms of Usage   Privacy Policy   Code of Ethics   Contact Us
Did you know the ACM DL App is now available?
Did you know your Organization can subscribe to the ACM Digital Library?
The ACM Guide to Computing Literature
All Tags
Export Formats
 
 
Save to Binder