Correlizer

From BOINC Projects
Revision as of 14:04, 30 August 2026 by Al Piskun (talk | contribs)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search



Correlizer
Project
StatusCompleted
CategoryBiology
ComputeCPU
RequiresNone
Development
DeveloperTobias A. Knoch
AuthorTobias A. Knoch
SponsorBioQuant Centre, University of Heidelberg (with the German Cancer Research Center and Erasmus MC)
MaintainerTobias A. Knoch
Initial releaseAugust 3, 2011  (15 years ago)
Completedc. 2015
Software
Operating systemWindows, Linux, macOS (also ARM/Linux embedded)
BOINC statistics
Stats as of2013-03-27
Performance1,160 GFLOPS
Active users615
Total users1,747
Active hosts2,096
Total hosts8,349
Analytics
GPU performanceNone (CPU only)
CPU performance1,160 GFLOPS
Metadata
Websitehttp://svahesrv2.bioquant.uni-heidelberg.de/correlizer/

Correlizer was a scientific volunteer computing project that ran on the BOINC platform and applied long-range correlation analysis to completely sequenced genomes in order to study the sequential and three-dimensional organization of genomes. It was operated from the Biophysical Genomics group at the BioQuant Centre of the University of Heidelberg in Germany, and the underlying science was developed by Tobias A. Knoch, a researcher also affiliated with the German Cancer Research Center (DKFZ) and the Erasmus MC in Rotterdam. The project began distributing work in August 2011 and is generally regarded as a completed BOINC project, having stopped issuing work units around 2015 after its workload was folded into AlmereGrid.[1][2]

Background: genome organization

The central motivation for Correlizer was a long-standing question in biology: how the informational content of a genome is related to the physical shape in which it is stored. In a human cell the genetic information is carried by a diploid set of DNA molecules, the chromosomes. The total sequence is about 3×109 base pairs (bp), which corresponds to roughly 2.80 GB of data. Laid end to end the molecule would be about 2 m long, yet it is folded into a cell nucleus with a typical diameter of about 10 μm and a volume of roughly 500 μm3.[3][4]

The sequential organization of a genome, meaning the statistical relations between base pairs that are far apart along the sequence, and the way that sequence relates to the three-dimensional architecture of the chromosomes, remained largely unresolved. Correlizer treated correlation analysis as a kind of "virtual microscope" for this hidden structure. The analysis is based on the concentration profile of single nucleotides along the DNA sequence. If cl is the concentration of a nucleotide in a window of length l, and the concentration in the whole sequence of length L is known, then the concentration-fluctuation function C(l) and the local correlation coefficient δ(l) can be computed. For a purely random sequence δ0.5, whereas values clearly above this threshold indicate positive long-range correlations; a power-law dependence such as C(l)lδ is the signature of correlations that extend over many length scales.[4]

What was computed

Correlizer distributed this correlation analysis over the volunteer computers of its participants. The application, delivered by the project as "Correlizer Applications" (version 1.09) and an embedded build for ARM/Linux (version 1.10), used the BOINC scheduler to farm out individual analysis jobs.[5]

A computer rendering of the DNA double helix
The genomic sequence analysed by Correlizer is carried by the DNA double helix.

The project found long-range power-law correlations on almost the entire observable scale of 132 completely sequenced chromosomes, ranging in length from 0.5×106 to 3.0×107 bp and spanning organisms as different as Archaea, Bacteria, Arabidopsis thaliana, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Drosophila melanogaster, and Homo sapiens. The local correlation coefficients displayed a species-specific multiscaling behavior: correlations were close to random on the scale of a few base pairs, rose to a first maximum between 40 and 3,400 bp (split into two submaxima for Arabidopsis thaliana and Drosophila melanogaster), and often showed one or more second maxima in the region from 105 to 3×105 bp.[4][1]

Within this multiscaling a further fine structure was present, attributable to codon usage in every species except the human sequences, where it instead reflected nucleosomal binding. The study also showed that computer-generated random sequences, designed to assume a block organization of genomes with a particular codon usage and nucleosomal binding, reproduced the observed behavior, while random reshuffling of real sequences destroyed the correlations. This suggested that the stability of the correlations is tightly controlled by evolution and closely connected to the spatial organization of the genome, particularly at large scales.[4]

Project history

Correlizer was announced to the BOINC community in August 2011 and hosted at svahesrv2.bioquant.uni-heidelberg.de/correlizer, with the server running on the BioQuant infrastructure at the University of Heidelberg.[2][6] The client application was CPU-based, with builds for Windows, Linux, macOS, and an embedded build for ARM-based Linux systems. The software could operate through a proxy and could be configured to run as a screensaver, but no GPU (CUDA or OpenCL) version was provided.[2][5]

A diagram of the human karyotype
A human karyotype. The 23 pairs of chromosomes carry roughly 3×109 base pairs of DNA.

By late 2014 the project had gone offline while the system and applications were updated, and the project's message boards noted that much of the older account and forum data had been lost. Volunteers reported over the following months that the project was no longer issuing work units, and by mid-2015 the project was widely considered inactive.[7][8] German-language project records list its end date as the point at which it stopped emitting work units because its workload had in the meantime been integrated into AlmereGrid, a Dutch volunteer computing project.[2] The project is consequently listed among the completed and retired BOINC projects.

Statistics at its last active snapshot

A server-status snapshot from 27 March 2013 recorded a healthy, running infrastructure. At that point the project reported about 416,000 tasks ready to send, roughly 103,500 tasks in progress, and an average work-unit runtime of about 0.26 hours. The volunteer base was small but steady, and the aggregate processing power reported was about 1,160 GFLOPS, i.e. roughly 1.16×1012 floating-point operations per second.[9]

Scientific publications

The following publications are credited to the project on the BOINC publications list.[10]

A fifth entry appears under the Correlizer heading on the BOINC list, "ISDEP, a fusion application deployed in the EDGeS project" by Alejandro Rivero and Dario Ferrer (2010). That entry concerns the EDGeS@Home project and appears to be a listing error on the aggregation page rather than a Correlizer result.[10][11]

Related publications by the project investigators

The computational framework used by Correlizer grew out of the same research program, and the project's methods were subsequently extended in further work by its principal investigator and collaborators:

See also

References

  1. 1.0 1.1 Correlizer: Revealing the Mysteries of Genome Organization. Biophysical Genomics. BioQuant Centre, University of Heidelberg. Retrieved 30 August 2026.
  2. 2.0 2.1 2.2 2.3 Correlizer (beendet). Rechenkraft. Retrieved 30 August 2026.
  3. Background: Approaching the three-Dimensional Organization of the Human Genome. Biophysical Genomics. BioQuant Centre, University of Heidelberg. Retrieved 30 August 2026.
  4. 4.0 4.1 4.2 4.3 4.4 4.5 (2009).Fine-structured multi-scaling long-range correlations in completely sequenced genomes: features, origin, and classification. European Biophysics Journal. pp. 757-779. DOI: 10.1007/s00249-009-0489-y.
  5. 5.0 5.1 Applications. Biophysical Genomics. BioQuant Centre, University of Heidelberg. Retrieved 30 August 2026.
  6. Correlizer Project (BOINC message boards). University of California, Berkeley. Retrieved 30 August 2026.
  7. Welcome back, Correlizer. University of California, Berkeley. Retrieved 30 August 2026.
  8. Projects to retire from your site stats that are no longer active. BOINC Combined Statistics. Retrieved 30 August 2026.
  9. Project status. Biophysical Genomics. BioQuant Centre, University of Heidelberg. Retrieved 30 August 2026.
  10. 10.0 10.1 Publications by BOINC Projects. University of California, Berkeley. Retrieved 30 August 2026.
  11. ISDEP, a fusion application deployed in the EDGeS project. Retrieved 30 August 2026.