SciLINC

Revision as of 18:29, 14 August 2026 by Al Piskun (talk | contribs) (first light)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)



SciLINC, an acronym for Scientific Literature Indexing on Networked Computers, was a BOINC volunteer computing project operated by the Missouri Botanical Garden in St. Louis, Missouri.[1] Volunteers attached their computers to the project so that pages of digitized botanical literature could be analyzed during idle time: each workunit bundled OCR text of scanned pages together with an algorithm for finding scientific names, and the results were used to build full text and keyword indexes for the Garden's Botanicus digital library.[2] SciLINC went online on March 22, 2007 and sent out its first workunits on March 29, 2007.[3] It was a pioneering effort, presented by its creators as the first attempt to apply public-resource computing within the library community.[1]

SciLINC
Project
StatusCompleted
CategoryBiology
ComputeCPU
RequiresBOINC client
Development
DeveloperMissouri Botanical Garden
AuthorMissouri Botanical Garden
SponsorInstitute of Museum and Library Services
Initial releaseMarch 22, 2007  (19 years ago)
CompletedJune 2007
Discontinued2009
Software
Operating systemMicrosoft Windows (98 or later), Linux
BOINC statistics
Stats as ofJuly 24, 2007  (19 years ago)
Performance0.0 GigaFLOPS
Active users0
Total users150
Active hosts0
Total hosts217
Analytics
RAC0
Credit/day0
Metadata
Websitehttps://www.scilinc.org/
LicenseNot published

Because text indexing demands very little computing per unit of data, the public phase of the project was brief: the available work was processed within days, and in June 2007 the team announced that all remaining SciLINC computation would be performed on the Garden's own servers.[4] Statistics sites such as BOINCstats froze their SciLINC data on July 24, 2007, recording 150 registered users, 217 hosts and 2,093 total credits.[5] The project website remained online through 2009 and later went offline.[3] Although short lived, SciLINC influenced later biodiversity informatics work, most notably the TaxonFinder name-finding service adopted by the Biodiversity Heritage Library.[2][6]

Background and goals

The Missouri Botanical Garden, founded in 1859, is one of the oldest botanical institutions in the United States and maintains a major research library of rare botanical literature.[7] Much of that literature is scarce or unique, and scientists have traditionally travelled to St. Louis to consult it. To make the collection widely available, the Garden digitizes books and journals through its Botanicus digital library, and in October 2005 it received a grant from the Institute of Museum and Library Services (IMLS), a United States federal agency, to build SciLINC as a three-year project.[1][8] Chris Freeland was the project's listed contact at the Garden, and Ron Parker wrote the project's news pages and development blog.[1][4][9]

The project description on Botanicus set out four primary goals:[1]

  1. Increase public access to nationally significant scientific literature held by the Missouri Botanical Garden Library.
  2. Enhance the usefulness of the digitized materials by creating a web repository of scanned literature, keywords and online resources with tools for searching and analysis.
  3. Create an educational tool for learning about plant life: while indexing keywords, the participant's computer would display information about each plant name being processed, including descriptive data, images, maps and annotated links.
  4. Provide a model for adopting public-resource computing applications within the library community, where such applications "have been successfully employed in research and museum communities" but "so far have not been applied within libraries".

The keywords identified by SciLINC were to be derived from the Garden's botanical databases and would include plant names, place names, authors and other phrases of interest to people studying plants, each annotated with links to related online resources.[1]

 
The Climatron conservatory at the Missouri Botanical Garden, which developed and operated SciLINC.

Design and operation

SciLINC used BOINC, the Berkeley Open Infrastructure for Network Computing, the open-source volunteer computing platform developed at the University of California, Berkeley, which also powers projects such as SETI@home and World Community Grid.[10] The project packaged OCR text of scanned botanical pages into workunits and sent them, along with a name-finding algorithm, to volunteers' computers, which then reported back the locations of scientific names and index terms found in the text.[2] Participants joined by installing the BOINC client and attaching to the project URL http://www.scilinc.org/SciLINC/.[4]

The volunteer application was named TaxonGrab. As listed on the project's applications page, version 1.39 ran on Microsoft Windows 98 or later on Intel x86-compatible processors, while the first results reported on March 29, 2007 were produced by a Linux version.[3][11] Participants operated under standard BOINC rules and policies: the software had to be run only on authorized computers, volunteers could control how much CPU power, disk space and bandwidth SciLINC used, and account names, along with a summary of each computer's work, could be shown on the project website.[12] Volunteers earned BOINC credit for completed work, though the team later acknowledged that "a relatively small amount of credit" was earned for "a comparatively large load" on participants' systems.[4]

During 2006, Parker documented development on the SciLINC Dev blog, covering topics such as installing the BOINC server software on CentOS and Fedora Core Linux, running the server under Xen virtualization, and a summary of a BOINC community conference question and answer session with BOINC founder David P. Anderson and SETI@home scientist Eric Korpela.[13][14] The blog also recorded the project's design philosophy, borrowed from the BOINC literature: to make efficient use of public-resource computing, the compute time of the client must be relatively large compared to the data-transfer time required, that is

tcomputettransfer

History

SciLINC went online on March 22, 2007, with the project team expecting to "be processing work within a few days".[3] On March 29, 2007 the server sent out its first set of workunits and received its first results, running on Linux, with a Windows version announced as coming shortly.[3] The project had originally operated on the Garden's internal network, and on May 14, 2007 the server was moved to the public Internet at www.scilinc.org; a server outage followed a few days later, and the recovery process, which involved changing the BOINC server's name, was documented on the development blog.[3][15]

On June 1, 2007 the project announced that results from SciLINC were being passed to Botanicus.[3] In mid June the project hit a series of setbacks: on June 15 the process that gathered work for SciLINC "filled the drive with everything it could find in the vaults" (images, PDFs and more); on June 16 an electrical transient in the building housing the server corrupted the MySQL database, including the SciLINC results table; and on June 18 new account creation was temporarily disabled.[3] Around the same time, volunteers on the BOINCstats and BOINC forums reported that the roughly 50 files served per workunit were causing high CPU load in the core BOINC client.[3] On June 19 the team restored the database but disabled the sending of new workunits, warning that pending results might never receive credit.[3]

The decisive announcement came on June 22, 2007 in a news item titled "SciLINC Update". The team reported that goals 1, 2 and 4 had been met (Botanicus was processing SciLINC's output, and the project had demonstrated the library use of public-resource computing), but that "the workunits fly by so rapidly" that the educational screensaver of goal 3 had never become realistic. It also noted that the BOINC community had discovered the project before the team was ready to announce it. The update concluded that "for now all SciLINC computation will be performed internally", with hopes of returning to volunteers with more computationally intensive analyses and "a better credit-reward ratio (and nicer screensaver)".[4] Internal testing continued at the end of June, with the scheduler briefly unavailable to Internet users on June 27, 2007.[4]

Public participation wound down quickly after that. BOINCstats, which had last imported SciLINC's statistics on July 24, 2007, later labelled the project "no longer active" with its statistics "frozen on the day of the last update".[5] The website, still bearing the June 2007 news, remained online through 2009, and was gone by the late 2010s.[3]

Data economics and the end of public computing

The June 22, 2007 update explained in unusual detail why text indexing was a poor fit for volunteer computing. The team calculated that SciLINC would have needed to transfer roughly 250 MiB of compressed data to occupy a modern CPU for one day; this would expand to nearly 660 MiB of input data, and the client would then upload about 44 MiB of results, which would compress to 17 MiB. The input data therefore expanded by a factor of

660/250=2.64

while the results compressed by a similar factor,

44/172.59

These numbers, the team said, "have only grown as SciLINC has been improved and made more efficient", and they were "not acceptable to the average BOINC user". The project lead's working rule, stated in the same update, was that from a technological and economic standpoint it makes sense to use public-resource computing in place of an internal grid architecture whenever less than a gigabyte of data is required per CPU-day of computation, that is

D<1 GiB per CPU-day

For a volunteer on dial-up giving SciLINC just 1% of their BOINC time, the daily computation would last

0.01×24×60=14.4

minutes, about 15 minutes, during which roughly 2.5 MiB would have to be downloaded; the team observed that for dial-up users "the transfer time would exceed the computation time". Even setting transfer concerns aside, the Garden simply did not have enough text to "realistically occupy hundred or thousands of BOINC enthusiasts for a lengthy period of time", and the work itself, "text-indexing and taxonomic analysis", had "a relatively low mathematical complexity". Freeland later summarized the experience: "we ran out of jobs in about 2 days because text indexing isn't processor intensive; the best BOINC projects take small inputs that require lots of computing resources."[4][2]

Project statistics

The following figures are the final values recorded by BOINCstats before the project's statistics were frozen on July 24, 2007:[5]

SciLINC statistics at BOINCstats (frozen)
Metric Value
Total users 150
Active users 0
Total hosts 217
Active hosts 0
Teams 35
Countries 23
Total credit 2,093
Recent average credit 0
Average performance 0.0 GigaFLOPS

Results and legacy

SciLINC's principal scientific output was the indexing of digitized botanical literature that fed the Garden's Botanicus portal, where full text searching and name-based discovery were integrated beginning in June 2007.[3] By December 2008 Botanicus reported 577 titles (books and journals), 2,548 volumes and 1,079,660 pages online, along with 184,750 links to protologues (original species descriptions).[1]

The project's longest lasting influence was methodological. Freeland wrote that SciLINC preceded the team's work on TaxonFinder, the algorithm and service, provided by uBio, that the Biodiversity Heritage Library (BHL) incorporated into its portal to identify and verify taxonomic name strings across its digitized corpus.[2][6] The experience also taught the team how jobs are packaged for distribution to volunteers "and the kinds of inputs and outputs that are successful in an asynchronous computing environment", a lesson Freeland later applied to proposals for "purposeful gaming" for the BHL, where human spare cycles, rather than CPU cycles, would improve digitized library metadata.[2] That idea grew into BHL's Purposeful Gaming project, funded by IMLS and run by the Garden's Center for Biodiversity Informatics from 2013 to 2015.[16] The team's lessons were collected in the unpublished IMLS final report, "SciLINC Final Report", with an appendix, linked from Freeland's 2011 blog post.[2]

 
A hand coloured plate from Curtis's Botanical Magazine, typical of the historical botanical literature that Botanicus digitizes and that SciLINC indexed.

Related publications

SciLINC itself did not publish scientific papers; its findings were reported to the funding agency and in blog posts. The techniques and systems it demonstrated appear in the following works. The first describes TaxonGrab, a rule-based technique for extracting taxonomic names from text; SciLINC's Windows volunteer application carried the same name.

The Anderson paper is included because it was the description of BOINC recommended by the SciLINC development blog as background reading for the project.[17]

See also

References

  1. 1.0 1.1 1.2 1.3 1.4 1.5 1.6 Botanicus.org: SciLINC, Scientific Literature Indexing on Networked Computers. Missouri Botanical Garden. Retrieved August 14, 2026.
  2. 2.0 2.1 2.2 2.3 2.4 2.5 2.6 Freeland, Chris.(December 2, 2011).BOINCing Angry Birds for BHL: Purposeful Gaming in Digital Libraries. ChrisFreeland.com. Retrieved August 14, 2026.
  3. 3.00 3.01 3.02 3.03 3.04 3.05 3.06 3.07 3.08 3.09 3.10 3.11 Parker, Ron.(June 19, 2007).SciLINC: News archive. Missouri Botanical Garden. Retrieved August 14, 2026.
  4. 4.0 4.1 4.2 4.3 4.4 4.5 4.6 Parker, Ron.(June 22, 2007).SciLINC: SciLINC Update. Missouri Botanical Garden. Retrieved August 14, 2026.
  5. 5.0 5.1 5.2 BOINCstats: SciLINC, Credit overview. BOINCstats. Retrieved August 14, 2026.
  6. 6.0 6.1 (October 17, 2008).Taxonomic Name Recognition in Biodiversity Heritage Library. University of Illinois at Urbana-Champaign, IDEALS. Retrieved August 14, 2026.
  7. Missouri Botanical Garden. Wikipedia. Retrieved August 14, 2026.
  8. Christopher David Freeland: Curriculum Vitae. ChrisFreeland.com. Retrieved August 14, 2026.
  9. Parker, Ron.(October 3, 2006).SciLINC, Scientific Literature Indexing on Networked Computers. SciLINC Dev. Retrieved August 14, 2026.
  10. BOINC: Open-source software for volunteer computing. University of California, Berkeley. Retrieved August 14, 2026.
  11. (June 27, 2007).SciLINC: Applications. Missouri Botanical Garden. Retrieved August 14, 2026.
  12. SciLINC: Read our rules and policies. Missouri Botanical Garden. Retrieved August 14, 2026.
  13. Parker, Ron.(November 19, 2006).Linux BOINC Server CentOS-4 Installation. SciLINC Dev. Retrieved August 14, 2026.
  14. Parker, Ron.(December 19, 2006).BOINC Conference Summary. SciLINC Dev. Retrieved August 14, 2026.
  15. Parker, Ron.(May 21, 2007).Problems changing a BOINC server's name. SciLINC Dev. Retrieved August 14, 2026.
  16. Purposeful Gaming. Biodiversity Heritage Library. Retrieved August 14, 2026.
  17. Parker, Ron.(October 3, 2006).SciLINC, Scientific Literature Indexing on Networked Computers. SciLINC Dev. Retrieved August 14, 2026.

External links