SciLINC: Difference between revisions

Al Piskun (talk | contribs)
first light
Tag: Blanking
Al Piskun (talk | contribs)
No edit summary
Tag: Manual revert
 
Line 1: Line 1:
{{Infobox software
| name                = SciLINC
| logo                =
| logo caption        =
| screenshot          =
| caption              = The BOINC Manager client software used by SciLINC volunteers
| description          = SciLINC (Scientific Literature Indexing on Networked Computers) was a short lived BOINC volunteer computing project run by the Missouri Botanical Garden that indexed digitized botanical literature on participants' computers between March and June 2007.
| status              = Completed
| category            = Biology
| compute              = CPU
| dependencies        = BOINC client
| developer            = Missouri Botanical Garden
| author              = Missouri Botanical Garden
| sponsor              = Institute of Museum and Library Services
| maintainer          =
| released            = {{Start date and age|2007|03|22}}
| completed            = June 2007
| discontinued        = 2009
| repository          =
| programming language =
| operating system    = Microsoft Windows (98 or later), Linux
| size                =
| stats as of          = {{Start date and age|2007|07|24}}
| average performance  = 0.0 GigaFLOPS
| active users        = 0
| total users          = 150
| active hosts        = 0
| total hosts          = 217
| rac                  = 0
| credit per day      = 0
| gpu performance      =
| cpu performance      =
| website              = {{URL|https://www.scilinc.org/}}
| license              = Not published
}}


'''SciLINC''', an acronym for '''Sci'''entific '''L'''iterature '''I'''ndexing on '''N'''etworked '''C'''omputers, was a [[BOINC]] volunteer computing project operated by the Missouri Botanical Garden in St. Louis, Missouri.<ref name="botanicus">{{cite web |url=http://www.botanicus.org/Scilinc.aspx |title=Botanicus.org: SciLINC, Scientific Literature Indexing on Networked Computers |publisher=Missouri Botanical Garden |access-date=August 14, 2026 |archive-url=https://web.archive.org/web/20081212014154/http://www.botanicus.org/Scilinc.aspx |archive-date=December 12, 2008 |url-status=dead}}</ref> Volunteers attached their computers to the project so that pages of digitized botanical literature could be analyzed during idle time: each workunit bundled OCR text of scanned pages together with an algorithm for finding scientific names, and the results were used to build full text and keyword indexes for the Garden's Botanicus digital library.<ref name="freeland2011">{{cite web |url=http://blog.chrisfreeland.com/2011/12/boincing-angry-birds-for-bhl.html |title=BOINCing Angry Birds for BHL: Purposeful Gaming in Digital Libraries |first=Chris |last=Freeland |publisher=ChrisFreeland.com |date=December 2, 2011 |access-date=August 14, 2026 |archive-url=https://web.archive.org/web/20120313135210/http://blog.chrisfreeland.com/2011/12/boincing-angry-birds-for-bhl.html |archive-date=March 13, 2012 |url-status=live}}</ref> SciLINC went online on March 22, 2007 and sent out its first workunits on March 29, 2007.<ref name="oldnews">{{cite web |url=http://www.scilinc.org/old_news.php |title=SciLINC: News archive |first=Ron |last=Parker |publisher=Missouri Botanical Garden |date=June 19, 2007 |access-date=August 14, 2026 |archive-url=https://web.archive.org/web/20071223012202/http://www.scilinc.org/old_news.php |archive-date=December 23, 2007 |url-status=dead}}</ref> It was a pioneering effort, presented by its creators as the first attempt to apply public-resource computing within the library community.<ref name="botanicus"/>
Because text indexing demands very little computing per unit of data, the public phase of the project was brief: the available work was processed within days, and in June 2007 the team announced that all remaining SciLINC computation would be performed on the Garden's own servers.<ref name="june22">{{cite web |url=http://www.scilinc.org/ |title=SciLINC: SciLINC Update |first=Ron |last=Parker |publisher=Missouri Botanical Garden |date=June 22, 2007 |access-date=August 14, 2026 |archive-url=https://web.archive.org/web/20071023063008/http://www.scilinc.org/ |archive-date=October 23, 2007 |url-status=dead}}</ref> Statistics sites such as [[BOINCstats]] froze their SciLINC data on July 24, 2007, recording 150 registered users, 217 hosts and 2,093 total credits.<ref name="boincstats">{{cite web |url=https://boincstats.com/stats/project_graph.php?pr=scilinc |title=BOINCstats: SciLINC, Credit overview |publisher=BOINCstats |access-date=August 14, 2026 |archive-url=https://web.archive.org/web/20111012020502/http://boincstats.com/stats/project_graph.php?pr=scilinc |archive-date=October 12, 2011 |url-status=dead}}</ref> The project website remained online through 2009 and later went offline.<ref name="oldnews"/> Although short lived, SciLINC influenced later biodiversity informatics work, most notably the TaxonFinder name-finding service adopted by the Biodiversity Heritage Library.<ref name="freeland2011"/><ref name="idealstnr">{{cite web |url=https://www.ideals.illinois.edu/handle/2142/9140 |title=Taxonomic Name Recognition in Biodiversity Heritage Library |first1=Qin |last1=Wei |first2=Chris |last2=Freeland |first3=P. Bryan |last3=Heidorn |publisher=University of Illinois at Urbana-Champaign, IDEALS |date=October 17, 2008 |access-date=August 14, 2026}}</ref>
== Background and goals ==
The Missouri Botanical Garden, founded in 1859, is one of the oldest botanical institutions in the United States and maintains a major research library of rare botanical literature.<ref>{{cite web |url=https://en.wikipedia.org/wiki/Missouri_Botanical_Garden |title=Missouri Botanical Garden |publisher=Wikipedia |access-date=August 14, 2026}}</ref> Much of that literature is scarce or unique, and scientists have traditionally travelled to St. Louis to consult it. To make the collection widely available, the Garden digitizes books and journals through its Botanicus digital library, and in October 2005 it received a grant from the Institute of Museum and Library Services (IMLS), a United States federal agency, to build SciLINC as a three-year project.<ref name="botanicus"/><ref>{{cite web |url=http://www.chrisfreeland.com/Freeland_CV.pdf |title=Christopher David Freeland: Curriculum Vitae |publisher=ChrisFreeland.com |access-date=August 14, 2026}}</ref> Chris Freeland was the project's listed contact at the Garden, and Ron Parker wrote the project's news pages and development blog.<ref name="botanicus"/><ref name="june22"/><ref>{{cite web |url=https://scilincdev.blogspot.com/2006/10/boinc.html |title=SciLINC, Scientific Literature Indexing on Networked Computers |first=Ron |last=Parker |work=SciLINC Dev |date=October 3, 2006 |access-date=August 14, 2026}}</ref>
The project description on Botanicus set out four primary goals:<ref name="botanicus"/>
# Increase public access to nationally significant scientific literature held by the Missouri Botanical Garden Library.
# Enhance the usefulness of the digitized materials by creating a web repository of scanned literature, keywords and online resources with tools for searching and analysis.
# Create an educational tool for learning about plant life: while indexing keywords, the participant's computer would display information about each plant name being processed, including descriptive data, images, maps and annotated links.
# Provide a model for adopting public-resource computing applications within the library community, where such applications "have been successfully employed in research and museum communities" but "so far have not been applied within libraries".
The keywords identified by SciLINC were to be derived from the Garden's botanical databases and would include plant names, place names, authors and other phrases of interest to people studying plants, each annotated with links to related online resources.<ref name="botanicus"/>
[[File:Climatron, Missouri Botanical Gardens.jpg|thumb|The Climatron conservatory at the Missouri Botanical Garden, which developed and operated SciLINC.]]
== Design and operation ==
SciLINC used [[BOINC]], the Berkeley Open Infrastructure for Network Computing, the open-source volunteer computing platform developed at the University of California, Berkeley, which also powers projects such as [[SETI@home]] and [[World Community Grid]].<ref>{{cite web |url=https://boinc.berkeley.edu/ |title=BOINC: Open-source software for volunteer computing |publisher=University of California, Berkeley |access-date=August 14, 2026}}</ref> The project packaged OCR text of scanned botanical pages into workunits and sent them, along with a name-finding algorithm, to volunteers' computers, which then reported back the locations of scientific names and index terms found in the text.<ref name="freeland2011"/> Participants joined by installing the BOINC client and attaching to the project URL <code>http://www.scilinc.org/SciLINC/</code>.<ref name="june22"/>
The volunteer application was named TaxonGrab. As listed on the project's applications page, version 1.39 ran on Microsoft Windows 98 or later on Intel x86-compatible processors, while the first results reported on March 29, 2007 were produced by a Linux version.<ref name="oldnews"/><ref>{{cite web |url=http://www.scilinc.org/apps.php |title=SciLINC: Applications |publisher=Missouri Botanical Garden |date=June 27, 2007 |access-date=August 14, 2026 |archive-url=https://web.archive.org/web/20070627230850/http://www.scilinc.org/apps.php |archive-date=June 27, 2007 |url-status=dead}}</ref> Participants operated under standard BOINC rules and policies: the software had to be run only on authorized computers, volunteers could control how much CPU power, disk space and bandwidth SciLINC used, and account names, along with a summary of each computer's work, could be shown on the project website.<ref>{{cite web |url=http://www.scilinc.org/info.php |title=SciLINC: Read our rules and policies |publisher=Missouri Botanical Garden |access-date=August 14, 2026 |archive-url=https://web.archive.org/web/20071023063008/http://www.scilinc.org/info.php |archive-date=October 23, 2007 |url-status=dead}}</ref> Volunteers earned BOINC credit for completed work, though the team later acknowledged that "a relatively small amount of credit" was earned for "a comparatively large load" on participants' systems.<ref name="june22"/>
During 2006, Parker documented development on the SciLINC Dev blog, covering topics such as installing the BOINC server software on CentOS and Fedora Core Linux, running the server under Xen virtualization, and a summary of a BOINC community conference question and answer session with BOINC founder David P. Anderson and SETI@home scientist Eric Korpela.<ref>{{cite web |url=https://scilincdev.blogspot.com/2006/11/linux-boinc-server-centos-4.html |title=Linux BOINC Server CentOS-4 Installation |first=Ron |last=Parker |work=SciLINC Dev |date=November 19, 2006 |access-date=August 14, 2026}}</ref><ref>{{cite web |url=https://scilincdev.blogspot.com/2006/12/boinc-conference-summary.html |title=BOINC Conference Summary |first=Ron |last=Parker |work=SciLINC Dev |date=December 19, 2006 |access-date=August 14, 2026}}</ref> The blog also recorded the project's design philosophy, borrowed from the BOINC literature: to make efficient use of public-resource computing, the compute time of the client must be relatively large compared to the data-transfer time required, that is
<math>t_{\text{compute}} \gg t_{\text{transfer}}</math>
== History ==
SciLINC went online on March 22, 2007, with the project team expecting to "be processing work within a few days".<ref name="oldnews"/> On March 29, 2007 the server sent out its first set of workunits and received its first results, running on Linux, with a Windows version announced as coming shortly.<ref name="oldnews"/> The project had originally operated on the Garden's internal network, and on May 14, 2007 the server was moved to the public Internet at www.scilinc.org; a server outage followed a few days later, and the recovery process, which involved changing the BOINC server's name, was documented on the development blog.<ref name="oldnews"/><ref>{{cite web |url=https://scilincdev.blogspot.com/2007/05/problems-changing-boinc-servers-name.html |title=Problems changing a BOINC server's name |first=Ron |last=Parker |work=SciLINC Dev |date=May 21, 2007 |access-date=August 14, 2026}}</ref>
On June 1, 2007 the project announced that results from SciLINC were being passed to Botanicus.<ref name="oldnews"/> In mid June the project hit a series of setbacks: on June 15 the process that gathered work for SciLINC "filled the drive with everything it could find in the vaults" (images, PDFs and more); on June 16 an electrical transient in the building housing the server corrupted the MySQL database, including the SciLINC results table; and on June 18 new account creation was temporarily disabled.<ref name="oldnews"/> Around the same time, volunteers on the BOINCstats and BOINC forums reported that the roughly 50 files served per workunit were causing high CPU load in the core BOINC client.<ref name="oldnews"/> On June 19 the team restored the database but disabled the sending of new workunits, warning that pending results might never receive credit.<ref name="oldnews"/>
The decisive announcement came on June 22, 2007 in a news item titled "SciLINC Update". The team reported that goals 1, 2 and 4 had been met (Botanicus was processing SciLINC's output, and the project had demonstrated the library use of public-resource computing), but that "the workunits fly by so rapidly" that the educational screensaver of goal 3 had never become realistic. It also noted that the BOINC community had discovered the project before the team was ready to announce it. The update concluded that "for now all SciLINC computation will be performed internally", with hopes of returning to volunteers with more computationally intensive analyses and "a better credit-reward ratio (and nicer screensaver)".<ref name="june22"/> Internal testing continued at the end of June, with the scheduler briefly unavailable to Internet users on June 27, 2007.<ref name="june22"/>
Public participation wound down quickly after that. BOINCstats, which had last imported SciLINC's statistics on July 24, 2007, later labelled the project "no longer active" with its statistics "frozen on the day of the last update".<ref name="boincstats"/> The website, still bearing the June 2007 news, remained online through 2009, and was gone by the late 2010s.<ref name="oldnews"/>
== Data economics and the end of public computing ==
The June 22, 2007 update explained in unusual detail why text indexing was a poor fit for volunteer computing. The team calculated that SciLINC would have needed to transfer roughly 250 MiB of compressed data to occupy a modern CPU for one day; this would expand to nearly 660 MiB of input data, and the client would then upload about 44 MiB of results, which would compress to 17 MiB. The input data therefore expanded by a factor of
<math>660/250 = 2.64</math>
while the results compressed by a similar factor,
<math>44/17 \approx 2.59</math>
These numbers, the team said, "have only grown as SciLINC has been improved and made more efficient", and they were "not acceptable to the average BOINC user". The project lead's working rule, stated in the same update, was that from a technological and economic standpoint it makes sense to use public-resource computing in place of an internal grid architecture whenever less than a gigabyte of data is required per CPU-day of computation, that is
<math>D < 1\ \text{GiB per CPU-day}</math>
For a volunteer on dial-up giving SciLINC just 1% of their BOINC time, the daily computation would last
<math>0.01 \times 24 \times 60 = 14.4</math>
minutes, about 15 minutes, during which roughly 2.5 MiB would have to be downloaded; the team observed that for dial-up users "the transfer time would exceed the computation time". Even setting transfer concerns aside, the Garden simply did not have enough text to "realistically occupy hundred or thousands of BOINC enthusiasts for a lengthy period of time", and the work itself, "text-indexing and taxonomic analysis", had "a relatively low mathematical complexity". Freeland later summarized the experience: "we ran out of jobs in about 2 days because text indexing isn't processor intensive; the best BOINC projects take small inputs that require lots of computing resources."<ref name="june22"/><ref name="freeland2011"/>
== Project statistics ==
The following figures are the final values recorded by BOINCstats before the project's statistics were frozen on July 24, 2007:<ref name="boincstats"/>
{| class="wikitable"
|+ SciLINC statistics at BOINCstats (frozen)
! Metric !! Value
|-
| Total users || 150
|-
| Active users || 0
|-
| Total hosts || 217
|-
| Active hosts || 0
|-
| Teams || 35
|-
| Countries || 23
|-
| Total credit || 2,093
|-
| Recent average credit || 0
|-
| Average performance || 0.0 GigaFLOPS
|}
== Results and legacy ==
SciLINC's principal scientific output was the indexing of digitized botanical literature that fed the Garden's Botanicus portal, where full text searching and name-based discovery were integrated beginning in June 2007.<ref name="oldnews"/> By December 2008 Botanicus reported 577 titles (books and journals), 2,548 volumes and 1,079,660 pages online, along with 184,750 links to protologues (original species descriptions).<ref name="botanicus"/>
The project's longest lasting influence was methodological. Freeland wrote that SciLINC preceded the team's work on TaxonFinder, the algorithm and service, provided by uBio, that the Biodiversity Heritage Library (BHL) incorporated into its portal to identify and verify taxonomic name strings across its digitized corpus.<ref name="freeland2011"/><ref name="idealstnr"/> The experience also taught the team how jobs are packaged for distribution to volunteers "and the kinds of inputs and outputs that are successful in an asynchronous computing environment", a lesson Freeland later applied to proposals for "purposeful gaming" for the BHL, where human spare cycles, rather than CPU cycles, would improve digitized library metadata.<ref name="freeland2011"/> That idea grew into BHL's Purposeful Gaming project, funded by IMLS and run by the Garden's Center for Biodiversity Informatics from 2013 to 2015.<ref>{{cite web |url=https://about.biodiversitylibrary.org/projects/past-projects/purposeful-gaming/ |title=Purposeful Gaming |publisher=Biodiversity Heritage Library |access-date=August 14, 2026}}</ref> The team's lessons were collected in the unpublished IMLS final report, "SciLINC Final Report", with an appendix, linked from Freeland's 2011 blog post.<ref name="freeland2011"/>
[[File:Curtis's botanical magazine (Plate 3709) (8386433065).jpg|thumb|upright=0.75|A hand coloured plate from ''Curtis's Botanical Magazine'', typical of the historical botanical literature that Botanicus digitizes and that SciLINC indexed.]]
== Related publications ==
SciLINC itself did not publish scientific papers; its findings were reported to the funding agency and in blog posts. The techniques and systems it demonstrated appear in the following works. The first describes TaxonGrab, a rule-based technique for extracting taxonomic names from text; SciLINC's Windows volunteer application carried the same name.
* {{cite journal |last1=Koning |first1=Drew |last2=Sarkar |first2=Indra Neil |last3=Moritz |first3=Thomas |year=2005 |title=TaxonGrab: Extracting taxonomic names from text |journal=Biodiversity Informatics |volume=2 |pages=79-82 |doi=10.17161/bi.v2i0.17}}
* {{cite web |url=http://precedings.nature.com/documents/3372/version/1 |title=An evaluation of taxonomic name finding & next steps in Biodiversity Heritage Library (BHL) developments |first=Chris |last=Freeland |publisher=Nature Precedings |year=2009 |access-date=August 14, 2026 |archive-url=https://web.archive.org/web/20090723202516/http://precedings.nature.com/documents/3372/version/1 |archive-date=July 23, 2009 |url-status=dead}}
* {{cite web |url=https://www.ideals.illinois.edu/handle/2142/9140 |title=Taxonomic Name Recognition in Biodiversity Heritage Library |first1=Qin |last1=Wei |first2=Chris |last2=Freeland |first3=P. Bryan |last3=Heidorn |publisher=University of Illinois at Urbana-Champaign, IDEALS |date=October 17, 2008 |access-date=August 14, 2026}}
* {{cite web |url=https://www.researchgate.net/publication/41492665 |title=Name Matters: Taxonomic Name Recognition (TNR) in Biodiversity Heritage Library (BHL) |first1=Qin |last1=Wei |first2=P. Bryan |last2=Heidorn |first3=Chris |last3=Freeland |publisher=ResearchGate |date=February 2010 |access-date=August 14, 2026}}
* {{cite web |url=https://boinc.berkeley.edu/grid_paper_04.pdf |title=BOINC: A System for Public-Resource Computing and Storage |first=David P. |last=Anderson |work=5th IEEE/ACM International Workshop on Grid Computing |date=November 8, 2004 |access-date=August 14, 2026}}
The Anderson paper is included because it was the description of BOINC recommended by the SciLINC development blog as background reading for the project.<ref>{{cite web |url=https://scilincdev.blogspot.com/2006/10/boinc.html |title=SciLINC, Scientific Literature Indexing on Networked Computers |first=Ron |last=Parker |work=SciLINC Dev |date=October 3, 2006 |access-date=August 14, 2026}}</ref>
== See also ==
* [[BOINC]]
* [[BOINC Manager]]
* [[BOINCstats]]
* [[SETI@home]]
* [[World Community Grid]]
== References ==
{{Reflist}}
== External links ==
* [http://www.botanicus.org/Scilinc.aspx SciLINC project description] at Botanicus.org
* [https://scilincdev.blogspot.com/ SciLINC Dev] development blog by Ron Parker
* [https://web.archive.org/web/20071023063008/http://www.scilinc.org/ SciLINC website] at the Wayback Machine
* [https://web.archive.org/web/20111012020502/http://boincstats.com/stats/project_graph.php?pr=scilinc SciLINC statistics] at BOINCstats via the Wayback Machine
[[Category:BOINC projects]]
[[Category:Completed projects]]