Digitala Vetenskapliga Arkivet

Ändra sökning
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Scalable Splitting of Massive Data Streams
Uppsala universitet, Teknisk-naturvetenskapliga vetenskapsområdet, Matematisk-datavetenskapliga sektionen, Institutionen för informationsteknologi, Datalogi. (UDBL)
Uppsala universitet, Teknisk-naturvetenskapliga vetenskapsområdet, Matematisk-datavetenskapliga sektionen, Institutionen för informationsteknologi, Datalogi. (UDBL)
2010 (Engelska)Ingår i: Database Systems for Advanced Applications: Part II / [ed] Kitagawa H., Ishikawa Y., Li Q., Watanabe C., Berlin: Springer-Verlag , 2010, s. 184-198Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Scalable execution of continuous queries over massive data streams often requires splitting input streams into parallel sub-streams over which query operators are executed in parallel. Automatic stream splitting is in general very difficult, as the optimal parallelization may depend on application semantics. To enable application specific stream splitting, we introduce splitstream functions where the user specifies non-procedural stream partitioning and replication. For high-volume streams, the stream splitting itself becomes a performance bottleneck. A cost model is introduced that estimates the performance of splitstream functions with respect to throughput and CPU usage. We implement parallel splitstream functions, and relate experimental results to cost model estimates. Based on the results, a splitstream function called autosplit is proposed, which scales well for high degrees of parallelism, and is robust for varying proportions of stream partitioning and replication. We show how user defined parallelization using autosplit provides substantially improved scalability (L = 64) over previously published results for the Linear Road Benchmark.

Ort, förlag, år, upplaga, sidor
Berlin: Springer-Verlag , 2010. s. 184-198
Serie
Lecture Notes in Computer Science, ISSN 0302-9743 ; 5982
Nyckelord [en]
distributed stream systems, parallelization, query optimization
Nationell ämneskategori
Datavetenskap (datalogi)
Identifikatorer
URN: urn:nbn:se:uu:diva-136403DOI: 10.1007/978-3-642-12098-5_15ISI: 000278934800015ISBN: 978-3-642-12097-8 (tryckt)OAI: oai:DiVA.org:uu-136403DiVA, id: diva2:376826
Konferens
15th International Conference, DASFAA 2010, Tsukuba, Japan, April 1-4, 2010
Projekt
iStreamseSSENCETillgänglig från: 2010-12-14 Skapad: 2010-12-13 Senast uppdaterad: 2018-01-12Bibliografiskt granskad
Ingår i avhandling
1. Scalable Parallelization of Expensive Continuous Queries over Massive Data Streams
Öppna denna publikation i ny flik eller fönster >>Scalable Parallelization of Expensive Continuous Queries over Massive Data Streams
2011 (Engelska)Doktorsavhandling, sammanläggning (Övrigt vetenskapligt)
Abstract [en]

Numerous applications in for example science, engineering, and financial analysis increasingly require online analysis over streaming data. These data streams are often of such a high rate that saving them to disk is not desirable or feasible. Therefore, search and analysis must be performed directly over the data in motion. Such on-line search and analysis can be expressed as continuous queries (CQs) that are defined over the streams. The result of a CQ is a stream itself, which is continuously updated as new data appears in the queried stream(s). In many cases, the applications require non-trivial analysis, leading to CQs involving expensive processing. To provide scalability of such expensive CQs over high-volume streams, the execution of the CQs must be parallelized.

In order to investigate different approaches to parallel execution of CQs, a parallel data stream management system called SCSQ was implemented for this Thesis. Data and queries from space physics and traffic management applications are used in the evaluations, as well as synthetic data and the standard data stream benchmark; the Linear Road Benchmark. Declarative parallelization functions are introduced into the query language of SCSQ, allowing the user to specify customized parallelization. In particular, declarative stream splitting functions are introduced, which split a stream into parallel sub-streams, over which expensive CQ operators are continuously executed in parallel.

Naïvely implemented, stream splitting becomes a bottleneck if the input streams are of high volume, if the CQ operators are massively parallelized, or if the stream splitting conditions are expensive. To eliminate this bottleneck, different approaches are investigated to automatically generate parallel execution plans for stream splitting functions. This Thesis shows that by parallelizing the stream splitting itself, expensive CQs can be processed at stream rates close to network speed. Furthermore, it is demonstrated how parallelized stream splitting allows orders of magnitude higher stream rates than any previously published results for the Linear Road Benchmark.

 

Ort, förlag, år, upplaga, sidor
Uppsala: Acta Universitatis Upsaliensis, 2011. s. 35
Serie
Digital Comprehensive Summaries of Uppsala Dissertations from the Faculty of Science and Technology, ISSN 1651-6214 ; 836
Nationell ämneskategori
Datavetenskap (datalogi)
Forskningsämne
Datavetenskap med inriktning mot databasteknik
Identifikatorer
urn:nbn:se:uu:diva-152255 (URN)978-91-554-8095-0 (ISBN)
Disputation
2011-09-20, Auditorium Minus, Museum Gustavianum, Akademigatan 3, Uppsala, 13:15 (Engelska)
Opponent
Handledare
Tillgänglig från: 2011-06-10 Skapad: 2011-04-27 Senast uppdaterad: 2018-01-12

Open Access i DiVA

fulltext(297 kB)1234 nedladdningar
Filinformation
Filnamn FULLTEXT02.pdfFilstorlek 297 kBChecksumma SHA-512
49e081e56f8e93ec70434feb24622c055b32809633c6eae2334cc5ae96685d780f06e3940ec5f903ff7c581ddebbd46faf2b1db9bad6f8b9482106eaae32204c
Typ fulltextMimetyp application/pdf

Övriga länkar

Förlagets fulltext

Sök vidare i DiVA

Av författaren/redaktören
Zeitler, ErikRisch, Tore
Av organisationen
Datalogi
Datavetenskap (datalogi)

Sök vidare utanför DiVA

GoogleGoogle Scholar
Totalt: 1240 nedladdningar
Antalet nedladdningar är summan av nedladdningar för alla fulltexter. Det kan inkludera t.ex tidigare versioner som nu inte längre är tillgängliga.

doi
isbn
urn-nbn

Altmetricpoäng

doi
isbn
urn-nbn
Totalt: 1003 träffar
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf