Opus, 2003 — X-Mine, NEC and the origins of our vector work
Primary source and corroborating filings for a claim in our stakeholder letter:
that we built early vector-embedding language models and named the product Opus.
The short version. In November 2003, NEC Corporation announced a bio-literature
mining product called Biocompass. The press release states it was licensed from
X-Mine, Inc. and built on the basic technology of Opus — which NEC defines as
"a text mining algorithm developed by Opus X-MINE that identifies direct and indirect
relationships between genes." That is a third party, more than two decades ago,
describing our approach in our own terms.
The primary source
NEC Corporation press release, 10 November 2003, "Launch of Biocompass, a Bio-related
Document Mining Tool." The release originally appeared at
nec.co.jp/press/ja/0311/1004.html; that URL no longer resolves, so the
archived copy is reproduced here in full.
NEC Corporation press release, 10 November 2003 (English rendering of the
Japanese original). Archived copy — the original NEC URL has since been retired.
"Biocompass is an APPLICATION LICENSE AGREEMENT (right to use software) with
X-MINE Inc., a U.S.-based bio-related software development company … and based on
the basic technology of Opus."
"(Note 2) A text mining algorithm developed by Opus X-MINE that identifies direct
and indirect relationships between genes."
Corroborating filings and publication
US2003/0204496A1 — "Inter-term relevance analysis for large libraries."
Sandip Ray, Raf Podowski, Kasian Franks. Assignee: X-Mine, Inc. Filed 29 April 2002.
View patent
US 7,987,191 — "System and method for generating a relationship network."
Kasian Franks, Cornelia A. Myers, Raf M. Podowski. Priority 6 June 2005, granted
26 July 2011. Claims a network of relationship vectors for what the filing calls
"hidden association and connection extraction."
View patent
Peer-reviewed publication — Nakazato T, Takinaka T, Mizuguchi H, Matsuda H,
Bono H, Asogawa M. "BioCompass: a novel functional inference tool that utilizes MeSH
hierarchy to analyze groups of genes." In Silico Biology 2008;8(1):53–61.
PMID 18430990
— NEC's own paper on the product. It documents the tool, not its licensing history.
Lawrence Berkeley National Laboratory — the same line of work was recognised
at Berkeley Lab's 2007 Excellence in Technology Transfer Awards
(Today at Berkeley Lab, 13 Dec 2007)
and won a
2008 R&D 100 Award
as the Biomimetic Search Engine.
Why it matters
Representing entities as vectors and reading the distances between them to surface
connections nobody had catalogued is the foundation of modern language models. We were
doing it on genes in 2002. We do it on markets now — the same mathematics, pointed at
a different corpus. The 2002 filing predates the 2013
word2vec paper that brought
the technique into the mainstream by more than a decade.
Published by Cymetica as a source record for our stakeholder
communications. Questions: contact@cymetica.com