{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T05:56:59Z","timestamp":1777442219981,"version":"3.51.4"},"reference-count":34,"publisher":"Wiley","issue":"9","license":[{"start":{"date-parts":[[2015,11,13]],"date-time":"2015-11-13T00:00:00Z","timestamp":1447372800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"funder":[{"name":"National Science Center","award":["DEC-2012\/07\/B\/ST6\/01516"],"award-info":[{"award-number":["DEC-2012\/07\/B\/ST6\/01516"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Concurrency and Computation"],"published-print":{"date-parts":[[2016,6,25]]},"abstract":"<jats:title>Summary<\/jats:title><jats:p>The paper presents a new open\u2010source framework called <jats:italic>KernelHive<\/jats:italic> for multilevel parallelization of computations among various clusters, cluster nodes, and finally, among both CPUs and GPUs for a particular application. An application is modeled as an acyclic directed graph with a possibility to run nodes in parallel and automatic expansion of nodes (called node unrolling) depending on the number of computation units available. A methodology is proposed for parallelization and mapping of an application to the environment that includes selection of devices using a chosen optimizer, selection of best grid configurations for compute devices, optimization of data partitioning and the execution. One of possibly many scheduling algorithms can be selected considering execution time, power consumption, and so on. An easy\u2010to\u2010use GUI is provided for modeling and monitoring with a repository of ready\u2010to\u2010use constructs and computational kernels. The methodology, execution times, and scalability have been demonstrated for a distributed and parallel password\u2010breaking example run in a heterogeneous environment with a cluster and servers with different numbers of nodes and both CPUs and GPUs. Additionally, performance of the framework has been compared with an MPI + OpenCL implementation using a parallel geospatial interpolation application employing up to 40 cluster nodes and 320 cores. Copyright \u00a9 2015 John Wiley &amp; Sons, Ltd.<\/jats:p>","DOI":"10.1002\/cpe.3719","type":"journal-article","created":{"date-parts":[[2015,11,14]],"date-time":"2015-11-14T00:48:37Z","timestamp":1447462117000},"page":"2586-2607","source":"Crossref","is-referenced-by-count":14,"title":["KernelHive: a new workflow\u2010based framework for multilevel high performance computing using clusters and workstations with CPUs and GPUs"],"prefix":"10.1002","volume":"28","author":[{"given":"Pawe\u0142","family":"Ro\u015bciszewski","sequence":"first","affiliation":[{"name":"Department of Computer Architecture, Faculty of Electronics, Telecommunications and Informatics Gdansk University of Technology  Narutowicza 11\/12 Gdansk 80\u2010233 Poland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pawe\u0142","family":"Czarnul","sequence":"additional","affiliation":[{"name":"Department of Computer Architecture, Faculty of Electronics, Telecommunications and Informatics Gdansk University of Technology  Narutowicza 11\/12 Gdansk 80\u2010233 Poland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rafa\u0142","family":"Lewandowski","sequence":"additional","affiliation":[{"name":"Department of Computer Architecture, Faculty of Electronics, Telecommunications and Informatics Gdansk University of Technology  Narutowicza 11\/12 Gdansk 80\u2010233 Poland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marcel","family":"Schally\u2010Kacprzak","sequence":"additional","affiliation":[{"name":"Department of Computer Architecture, Faculty of Electronics, Telecommunications and Informatics Gdansk University of Technology  Narutowicza 11\/12 Gdansk 80\u2010233 Poland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2015,11,13]]},"reference":[{"key":"e_1_2_10_2_1","volume-title":"Sourcebook of Parallel Computing","author":"Dongarra J","year":"2003"},{"key":"e_1_2_10_3_1","volume-title":"Parallel Programming: Techniques and Applications Using Networked Workstations and Parallel Computers","author":"Wilkinson B","year":"1999"},{"key":"e_1_2_10_4_1","volume-title":"Programming Massively Parallel Processors: A Hands\u2010on Approach","author":"Kirk DB","year":"2010"},{"key":"e_1_2_10_5_1","volume-title":"Foundations of Multithreaded, Parallel, and Distributed Programming","author":"Andrews GR","year":"2000"},{"key":"e_1_2_10_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0743-7315(03)00002-9"},{"key":"e_1_2_10_7_1","unstructured":"KellerR M\u00fcllerM.The grid\u2010computing library PACX\u2010MPI: extending MPI for computational grids 2012.http:\/\/www.hlrs.de\/organization\/av\/spmt\/research\/pacx-mpi\/[Accessed on 17 August 2015]."},{"key":"e_1_2_10_8_1","volume-title":"Using MPI: Portable Parallel Programming with the Message Passing Interface","author":"Gropp W","year":"1999"},{"key":"e_1_2_10_9_1","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/7055.001.0001","volume-title":"Using MPI\u20102: Advanced Features of the Message\u2010Passing Interface","author":"Gropp W","year":"1999"},{"key":"e_1_2_10_10_1","volume-title":"Proceedings of Parallel Processing and Applied Mathematics 2007 Conference","author":"Czarnul P","year":"2008"},{"key":"e_1_2_10_11_1","volume-title":"Multithreaded Programming with Pthreads","author":"Lewis B","year":"1998"},{"key":"e_1_2_10_12_1","volume-title":"Java Threads","author":"Oaks S","year":"2004"},{"key":"e_1_2_10_13_1","unstructured":"Khronos OpenCL Working Group.The OpenCL specification version 1.1 2011.http:\/\/www.khronos.org\/registry\/cl\/specs\/opencl-1.1.pdf[Accessed on 17 August 2015]."},{"key":"e_1_2_10_14_1","unstructured":"HeYH DingC.Hybrid OpenMP and MPI: programming and tuning. NUG2004 Lawrence Berkeley National Laboratory 2004."},{"key":"e_1_2_10_15_1","unstructured":"CAPS Enterprise Cray Inc. NVIDIA and the Portland Group.The openacc application programming interface v1.0 2011."},{"key":"e_1_2_10_16_1","series-title":"Euro\u2010Par'12","first-page":"859","volume-title":"Proceedings of the 18th International Conference on Parallel Processing","author":"Wienke S","year":"2012"},{"key":"e_1_2_10_17_1","unstructured":"BarakA ShilohA.The mosix virtual opencl (vcl) cluster platform 2011."},{"key":"e_1_2_10_18_1","doi-asserted-by":"crossref","unstructured":"BarakA ShilohELA Ben\u2010nunT.A package for OpenCL based heterogeneous computing on clusters with many gpu devices. InProceedings of International Conference on Cluster Computing:Heraklion Crete 2011;1\u20137.","DOI":"10.1109\/CLUSTERWKSP.2010.5613086"},{"issue":"2","key":"e_1_2_10_19_1","first-page":"26","article-title":"Distributed OpenCL distributing OpenCL platform on network scale","author":"Eskikaya B","year":"2012","journal-title":"IJCA Special Issue on Advanced Computing and Communication Technologies for HPC Applications"},{"key":"e_1_2_10_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-14122-5_44"},{"key":"e_1_2_10_21_1","doi-asserted-by":"crossref","unstructured":"DuatoJ PenaAJ SillaF MayoR Quintana\u2010OrtiES.RCUDA: reducing the number of GPU\u2010based accelerators in high performance clusters. In2010 International Conference on High Performance Computing and Simulation (HPCS) 2010;224\u2013231.","DOI":"10.1109\/HPCS.2010.5547126"},{"key":"e_1_2_10_22_1","unstructured":"HeC DuP.CUDA performance study on Hadoop MapReduce clusters 2010.http:\/\/www.slideshare.net\/ airbots\/cuda-29330283 University of Nebraska\u2010Lincoln [Accessed on 17 August 2015]."},{"key":"e_1_2_10_23_1","series-title":"Euro\u2010Par '09","first-page":"887","volume-title":"Proceedings of the 15th International Euro\u2010Par Conference on Parallel Processing","author":"Yan Y","year":"2009"},{"key":"e_1_2_10_24_1","unstructured":"TsiomenkoR ReesBS.Accelerating fast Fourier transforms using Hadoop and CUDA 2013.http: \/\/arxiv.org\/pdf\/1407.6915v1[Accessed on 17 August 2015]."},{"key":"e_1_2_10_25_1","unstructured":"LinY OkurS RadoiC.Hadoop+aparapi: making heterogenous MapReduce programming easier 2012. University of Illinois at Urbana Champaign."},{"key":"e_1_2_10_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-013-0912-0"},{"key":"e_1_2_10_27_1","doi-asserted-by":"crossref","unstructured":"AllardJ GourantonV LecointreL LimetS MelinE RaffinB RobertS.Flowvr: a middleware for large scale virtual reality applications. InProceedings of Euro\u2010Par 2004:Pisa Italia 2004;497\u2013505.","DOI":"10.1007\/978-3-540-27866-5_65"},{"key":"e_1_2_10_28_1","first-page":"277","volume-title":"14th IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing","author":"Dreher M","year":"2014"},{"key":"e_1_2_10_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-45249-9_5"},{"key":"e_1_2_10_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-012-0837-z"},{"key":"e_1_2_10_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-010-0499-7"},{"key":"e_1_2_10_32_1","unstructured":"BultrowiczS CzarnulP Ro\u015bciszewskiP.Runtime visualization of application progress and monitoring of a GPU\u2010enabled parallel environment. InProceedings of 13th International Conference on Software Engineering Parallel and Distributed Systems (SEPADS '14):Gdansk Poland 2014;70\u201379.http:\/\/www.wseas.us\/e-library\/conferences\/ 2014\/Gdansk\/SEBIO\/SEBIO-07.pdf[Accessed on 17 August 2015]."},{"key":"e_1_2_10_33_1","doi-asserted-by":"publisher","DOI":"10.5121\/ijcnc.2014.6506"},{"key":"e_1_2_10_34_1","series-title":"Computer Science & Information Technology","first-page":"9","volume-title":"ICCSEA2014","author":"Ro\u015bciszewski Pawe\u0142","year":"2014"},{"key":"e_1_2_10_35_1","volume-title":"Statistics for Spatial Data","author":"Cressie N","year":"2015"}],"container-title":["Concurrency and Computation: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/api.wiley.com\/onlinelibrary\/tdm\/v1\/articles\/10.1002%2Fcpe.3719","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.3719","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,2]],"date-time":"2023-09-02T22:01:09Z","timestamp":1693692069000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/cpe.3719"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,11,13]]},"references-count":34,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2016,6,25]]}},"alternative-id":["10.1002\/cpe.3719"],"URL":"https:\/\/doi.org\/10.1002\/cpe.3719","archive":["Portico"],"relation":{},"ISSN":["1532-0626","1532-0634"],"issn-type":[{"value":"1532-0626","type":"print"},{"value":"1532-0634","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,11,13]]}}}