Digitala Vetenskapliga Arkivet

Endre søk
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Processing High-Volume Stream Queries on a Supercomputer
Uppsala universitet, Teknisk-naturvetenskapliga vetenskapsområdet, Matematisk-datavetenskapliga sektionen, Institutionen för informationsteknologi. Teknisk-naturvetenskapliga fakulteten, Biologiska sektionen, Institutionen för ekologi och evolution, Datalogi. CSD. (UDBL)
Uppsala universitet, Teknisk-naturvetenskapliga vetenskapsområdet, Matematisk-datavetenskapliga sektionen, Institutionen för informationsteknologi. Teknisk-naturvetenskapliga fakulteten, Biologiska sektionen, Institutionen för ekologi och evolution, Datalogi. CSD. (UDBL)
2006 (engelsk)Inngår i: Processing High-Volume Stream Queries on a Supercomputer, 2006, s. 147-Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Scientific instruments, such as radio telescopes, colliders, sensor networks, and simulators generate very high volumes of data streams that scientists analyze to detect and understand physical phenomena. The high data volume and the need for advanced computations on the streams require substantial hardware resources and scalable stream processing. We address these challenges by developing data stream management technology to support high-volume stream queries utilizing massively parallel computer hardware. We have developed a data stream management system prototype for state-of-the-art parallel hardware. The performance evaluation uses real measurement data from LOFAR, a radio telescope antenna array being developed in the Netherlands.

sted, utgiver, år, opplag, sider
2006. s. 147-
HSV kategori
Identifikatorer
URN: urn:nbn:se:uu:diva-20450ISBN: 0-7695-2571-7 (tryckt)OAI: oai:DiVA.org:uu-20450DiVA, id: diva2:48223
Tilgjengelig fra: 2006-12-22 Laget: 2006-12-22 Sist oppdatert: 2018-01-12
Inngår i avhandling
1. Scalable Parallelization of Expensive Continuous Queries over Massive Data Streams
Åpne denne publikasjonen i ny fane eller vindu >>Scalable Parallelization of Expensive Continuous Queries over Massive Data Streams
2011 (engelsk)Doktoravhandling, med artikler (Annet vitenskapelig)
Abstract [en]

Numerous applications in for example science, engineering, and financial analysis increasingly require online analysis over streaming data. These data streams are often of such a high rate that saving them to disk is not desirable or feasible. Therefore, search and analysis must be performed directly over the data in motion. Such on-line search and analysis can be expressed as continuous queries (CQs) that are defined over the streams. The result of a CQ is a stream itself, which is continuously updated as new data appears in the queried stream(s). In many cases, the applications require non-trivial analysis, leading to CQs involving expensive processing. To provide scalability of such expensive CQs over high-volume streams, the execution of the CQs must be parallelized.

In order to investigate different approaches to parallel execution of CQs, a parallel data stream management system called SCSQ was implemented for this Thesis. Data and queries from space physics and traffic management applications are used in the evaluations, as well as synthetic data and the standard data stream benchmark; the Linear Road Benchmark. Declarative parallelization functions are introduced into the query language of SCSQ, allowing the user to specify customized parallelization. In particular, declarative stream splitting functions are introduced, which split a stream into parallel sub-streams, over which expensive CQ operators are continuously executed in parallel.

Naïvely implemented, stream splitting becomes a bottleneck if the input streams are of high volume, if the CQ operators are massively parallelized, or if the stream splitting conditions are expensive. To eliminate this bottleneck, different approaches are investigated to automatically generate parallel execution plans for stream splitting functions. This Thesis shows that by parallelizing the stream splitting itself, expensive CQs can be processed at stream rates close to network speed. Furthermore, it is demonstrated how parallelized stream splitting allows orders of magnitude higher stream rates than any previously published results for the Linear Road Benchmark.

 

sted, utgiver, år, opplag, sider
Uppsala: Acta Universitatis Upsaliensis, 2011. s. 35
Serie
Digital Comprehensive Summaries of Uppsala Dissertations from the Faculty of Science and Technology, ISSN 1651-6214 ; 836
HSV kategori
Forskningsprogram
Datavetenskap med inriktning mot databasteknik
Identifikatorer
urn:nbn:se:uu:diva-152255 (URN)978-91-554-8095-0 (ISBN)
Disputas
2011-09-20, Auditorium Minus, Museum Gustavianum, Akademigatan 3, Uppsala, 13:15 (engelsk)
Opponent
Veileder
Tilgjengelig fra: 2011-06-10 Laget: 2011-04-27 Sist oppdatert: 2018-01-12

Open Access i DiVA

Fulltekst mangler i DiVA

Andre lenker

http://user.it.uu.se/~torer/publ/zricde2006.pdf

Søk i DiVA

Av forfatter/redaktør
Zeitler, ErikRisch, Tore
Av organisasjonen

Søk utenfor DiVA

GoogleGoogle Scholar

isbn
urn-nbn

Altmetric

isbn
urn-nbn
Totalt: 638 treff
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf