Weixiang Hong et al. prototype-guided transfer of sparse literature knowledge for electrolyte additive discovery in Guangzhou, China; enrichment factors 9.2-45.6 while screening <2%

1

Prototype-guided transfer of sparse literature knowledge for electrolyte additive discovery

Weixiang HONG1,5, Hongting DU1,5, Jiayue TANG1, Ruifeng TAN1, Yangjian QUAN2, Jia LI3,, Jiaqiang HUANG1,4,

1 Sustainable Energy and Environment Thrust and Guangzhou Municipal Key Laboratory of Materials Informatics, The Hong Kong University of Science and Technology (Guangzhou), Nansha, Guangzhou, 511400, Guangdong, P.R. China 2 Department of Chemistry, The Hong Kong University of Science and Technology, Clear Water Bay, Kowloon, Hong Kong SAR, P.R. China 3 Data Science and Analytics Thrust and Guangzhou Municipal Key Laboratory of Materials Informatics, The Hong Kong University of Science and Technology (Guangzhou), Nansha, Guangzhou, 511400, Guangdong, P.R. China 4 Academy of Interdisciplinary Studies, The Hong Kong University of Science and Technology, Clear Water Bay, Kowloon, Hong Kong SAR, P.R. China 5 These authors contributed equally: Weixiang HONG, Hongting DU

2

Abstract Electrolyte additive discovery remains challenging because experimentally validated molecules are sparse, whereas accessible chemical spaces are vast and largely unlabeled. This challenge is amplified in lithium-ion batteries, where additive performance arises from coupled interfacial reactions rather than a single molecular property. Here, we develop a prototype-guided molecular intelligence, ProtoMI, a literature-driven framework that learns transferable structural priors from reported electrolyte additives and uses them to prioritize candidates in unlabeled chemical space. For boron-containing additives, ProtoMI combines 126 literature-reported molecules with 179,977 unlabeled candidates. Graph contrastive learning identifies seven chemically interpretable prototypes from the reported additives, and prototype- guided semi-supervised contrastive learning adapts these prototypes to the candidate space under source-target distribution mismatch. In retrospective temporal validation, ProtoMI achieves enrichment factors of 9.2-45.6 while screening less than 2% of the candidate space. A subsequent translation step identifies four commercially accessible candidates. One representative candidate, 4,4,5,5-Tetramethyl-2-[10-(1- naphthyl)anthracen-9-yl]-1,3,2-dioxaborolane (TNDB), improves high-temperature LiFePO4||graphite cycling at 55 °C by 34.93% relative to the baseline electrolyte. An arsenal of characterizations and operando optical fiber Fourier transform infrared spectroscopy suggest that TNDB forms B-containing, F/P/O-modified inorganic interphases, suppresses solvent decomposition and reduces Fe deposition on graphite. This case study shows how sparse literature knowledge can guide experimentally efficient molecular discovery in data-scarce battery-additive spaces.

Introduction Lithium-ion batteries (LIBs) underpin electrified transport, grid storage and portable electronics, and their deployment is expected to expand substantially as electrification accelerates1,2. As LIBs are pushed toward longer lifetimes, higher power operation and harsher operating environments, battery degradation is increasingly governed by the stability of electrochemical interfaces, where electrolyte decomposition, interphase evolution and transition-metal crossover can trigger coupled degradation processes3- 5. Electrolyte additives provide a practical and manufacturable route to regulate such interfaces. Even at low concentrations, additives can alter solvation structures, undergo preferential interfacial reactions, scavenge reactive impurities and promote stabilizing solid-electrolyte interphase (SEI) and cathode-electrolyte interphase (CEI)6. However, additive discovery remains largely empirical because additive efficacy emerges from the coupled effects of molecular stability, salt-solvent coordination, ion

3

transport, decomposition pathways and interphase growth, none of which can be reliably captured by a single molecular descriptor3-6. Recent advances in artificial intelligence (AI), molecular representation learning and data-driven optimization have accelerated the exploration of high-dimensional chemical spaces7,8. In battery-electrolyte research, data-driven strategies have been used9 to predict molecular and electrolyte properties10, screen electrolyte components11 with theoretical calculations12-16 or curated datasets17,18, and guide electrolyte optimization workflows18-20. Active-learning approaches have further reduced experimental burden by iteratively updating models with new measurements20, while physics-informed and solvation-aware models have introduced thermodynamic, transport and molecular-interaction constraints into electrolyte design21,22. These studies have demonstrated the value of data-driven electrolyte discovery. Nevertheless, most existing approaches are formulated around supervised prediction or iterative optimization once target labels, simulation outputs or experimental feedback are available. This creates a mismatch for electrolyte additive discovery, where experimentally validated molecules are sparse, literature reports are heterogeneous and most accessible candidate molecules have no measured battery- relevant labels. Under such severe label scarcity, the central challenge shifts from predicting the performance of individual molecules to prioritizing chemically promising regions of an unlabeled molecular space under limited experimental budgets. Reported electrolyte additives are not merely isolated successful examples. Instead, they may encode recurring structural organization and shared interfacial regulation strategies6, including anion reception23, reactive-species scavenging24, solvation regulation25 and interphase construction24,26. If these recurring patterns can be learned as prototype- level priors rather than as molecule-level labels, they may guide exploration beyond the local neighborhoods of known additives. Representation learning provides a route to extract structured molecular embeddings from sparse data27, while prototype-based alignment28,29 and semi-supervised learning30 offer mechanisms to transfer such organization into unlabeled domains. Together, these advances suggest an alternative discovery paradigm: instead of directly predicting additive performance from scarce labels, sparse literature-derived success patterns can be converted into transferable molecular prototypes that guide efficient candidate selection (Fig. 1a,b). Here we develop prototype-guided molecular intelligence (ProtoMI), a literature-

4

driven framework for molecular discovery under sparse positive knowledge and large unlabeled candidate spaces. ProtoMI treats reported additives as a positive molecular knowledge source rather than as a conventional supervised training set. The framework proceeds through three stages (Fig. 1c). First, literature-reported additive molecules are extracted and curated to construct a positive molecular dataset, and the unsupervised graph contrastive learning is used to identify prototype-level chemical organization within the reported additives. Second, prototype-guided semi-supervised contrastive learning adapts these prototypes to a large unlabeled molecular space under source-target distribution mismatch. Third, a subsequent knowledge-guided translation step incorporates chemical feasibility and experimental accessibility constraints to convert model recommendations into testable candidates. As a case study, we focus on boron-containing electrolyte additives. Boron- containing additives provide a suitable testbed because their Lewis-acidic and borate/boronate chemistries are well documented in electrolyte interphase regulation, yet remain sparsely explored across the broader boron-containing chemical space23- 26. This combination of established chemical relevance and large unexplored molecular diversity creates a stringent test for ProtoMI: whether sparse literature- reported successes can be transformed into transferable prototype priors for efficient candidate selection. In this work, ProtoMI integrates 126 literature-reported boron- containing additives with 179,977 unlabeled boron-containing candidates from PubChem31. Retrospective temporal validation shows that the framework recovers future-reported additives with enrichment factors of 9.2-45.6 while screening less than 2% of the candidate space. A knowledge-guided translation step further prioritizes four commercially accessible candidates for high-temperature LiFePO4||graphite cells, a system in which elevated-temperature cycling stability remains an important practical challenge32,33. One representative candidate, 4,4,5,5-tetramethyl-2-[10-(1- naphthyl)anthracen-9-yl]-1,3,2-dioxaborolane (TNDB), improves capacity retention by 34.93% relative to the baseline electrolyte during cycling at 55 °C. Post-mortem and operando characterizations indicate that TNDB participates in forming B-containing, F/P/O-modified interphases that suppress solvent decomposition and reduce Fe deposition on graphite. These results demonstrate that sparse literature-derived molecular successes can be transformed into transferable prototype priors for efficient electrolyte additive discovery under severe data scarcity.

5

Fig. 1 | Overview of the ProtoMI framework for prototype-guided discovery of boron-containing electrolyte additives. a,b, Conceptual comparison between traditional molecular exploration and prototype-guided exploration. Conventional exploration (a) searches the molecular space largely through local trial-and-error optimization, whereas prototype-guided exploration (b) first identifies representative prototype regions from reported additives and subsequently explores structurally distinct regions of the molecular landscape through prototype-level knowledge transfer. c, Overview of the ProtoMI molecular discovery workflow. Experimentally reported boron-containing additives are first collected from the literature using a large-language-model-assisted data-mining pipeline to construct a curated molecular dataset. An unsupervised graph contrastive learning model then extracts prototype- level chemical knowledge from reported additives. These prototypes are transferred to a large unlabeled molecular space by prototype-guided semi-supervised contrastive learning, yielding adapted prototype regions and recommended candidates. A knowledge-guided translation step further incorporates chemical feasibility and practical accessibility constraints to generate experimentally actionable candidates for validation.

Results Dataset construction reveals source-target mismatch in boron-additive chemical space We first formulated boron-containing electrolyte additive discovery as a sparse- positive molecular exploration problem. To define the target search space, we collected 179,977 boron-containing candidate molecules from PubChem31. In parallel, we constructed a positive dataset of 126 experimentally reported boron-containing electrolyte additives extracted from Web of Science articles using a large language model (LLM)-assisted heuristic pipeline (Fig. 2a and Methods M1). This design reflects a realistic discovery setting: successful additives are sparsely reported in the literature, whereas the accessible boron-containing chemical space is orders of magnitude larger

6

and remains largely unlabeled. To evaluate whether unstructured literature could provide reliable positive molecular knowledge, we manually verified the extracted information across multiple battery-relevant fields. The extraction pipeline achieved high accuracy for chemically important categories, including 94% accuracy for additive identification, while requiring approximately