physics.chem-phJun 18, 2026

Empowering Polymeric Materials Discovery by Artificial Intelligence

Authors: Chenyao MaLinda ZhangYuheng ChenWei DuShangwen FangZihao JiangChuanyu LiuXinyu Ma+24 more

Organizations: Suzhou MatSource Technology Co., Ltd., Suzhou 215000, Jiangsu, China. · Advanced Institute for Materials Research (WPI-AIMR), Tohoku University, Sendai 980-8577, Japan. · Frontier Research Institute for Interdisciplinary Sciences (FRIS), Tohoku University, Sendai, 980-8577, Japan. · Jiangsu Key Laboratory of New Power Batteries, Jiangsu Collaborative Innovation Centre of Biomedical Functional Materials, School of Chemistry and Materials Science, Nanjing Normal University, Nanjing 210023, P.R. China. · Gusu Laboratory of Materials, Suzhou 215000, Jiangsu, China. · Department of Chemistry, National University of Singapore, Singapore, Singapore. · Thrust of Sustainable Energy and Environment, The Hong Kong University of Science and Technology (Guangzhou), Guangdong, Guangzhou, 511453, China. · Department of Materials Design and Innovation, University at Buffalo, Buffalo, NY 14260, USA. · College of Smart Materials and Future Energy, State Key Laboratory of Molecular Engineering of Polymers, Fudan University, Shanghai 200433, China. · School of Physical Science and Technology, Shanghai tech University, Shanghai 201210, P.R. China. · The State Key Laboratory of Molecular Engineering of Polymers and Department of Macromolecular Science, Fudan University, Shanghai 200438, People’s Republic of China. · Department of Chemistry and Materials Science, Xi'an J Liverpool University, Suzhou 215123, Jiangsu, P. R. China. · Key Laboratory of Electroanalytical Chemistry, Changchun Institute of Applied Chemistry, Chinese Academy of Sciences, Changchun, 130022, China. · State Key Laboratory of Advanced Environmental Technology, Department of Environmental Science and Engineering, University of Science and Technology of China, 230026, China.

Abstract

Polymeric materials underpin modern technologies spanning energy storage, microelectronics, healthcare and sustainable manufacturing. Yet their rational design remains exceptionally challenging because material performance emerges from complex interactions among molecular composition, chain architecture, processing history and hierarchical structural evolution across multiple length and time scales. Consequently, polymer research has long relied on labor-intensive experimentation and fragmented modeling approaches, limiting both mechanistic understanding and innovation efficiency. Recent advances in data infrastructure, machine learning, large artificial intelligence (AI) models and laboratory automation are beginning to reshape this landscape. Rather than functioning as isolated tools, polymer databases, predictive models, AI agents and automated laboratories are increasingly converging into interconnected discovery ecosystems. As a result, the central challenge is shifting from improving predictive accuracy alone to enabling reliable decision-making, adaptive learning and seamless integration across computation, experimentation and scientific reasoning. We argue that polymer science is entering an era of autonomous discovery, in which data, simulation, reasoning and experimentation operate within self-improving feedback loops that continuously generate hypotheses, design materials, execute experiments and refine predictive models. By unifying molecular design, process optimization, experimental validation and industrial translation, such autonomous ecosystems establish a more predictive, reproducible and scalable paradigm for polymer innovation, fundamentally transforming how polymer research is conducted.

Explore similar work

May 26, 2026cs.AI

PolyFusionAgent: A Multimodal Foundation Model and Autonomous AI Assistant for Polymer Property Prediction and Inverse Design

Polymer discovery is central to fields ranging from energy storage to biomedicine, but it is hindered by an astronomically large chemical design space and fragmented representations of structure, properties, and prior knowledge. This fragmentation leaves many AI models disconnected from physical and experimental reality, restricting their ability to support directly actionable design decisions. Here we introduce PolyFusionAgent, an interactive framework coupling a multimodal polymer foundation model (PolyFusion) with a tool-augmented, literature-grounded design agent (PolyAgent). PolyFusion aligns complementary polymer views including sequence, topology, 3D geometry, and fingerprints across millions of polymers to learn a shared latent space transferable across chemistries and data regimes, improving thermophysical property prediction and enabling property-conditioned generation of chemically valid, structurally novel polymers beyond the reference design space. PolyAgent closes the design loop by linking prediction and inverse design with evidence retrieval from the polymer literature, proposing, evaluating, and contextualizing hypotheses with explicit precedent in one workflow. Together, PolyFusionAgent enables interactive, evidence-linked polymer discovery combining large-scale representation learning, multimodal chemical knowledge, and verifiable scientific reasoning.
Manpreet Kaur, Xingying Zhang, Qian Liu
May 7, 2026cs.LG

Can LLMs Predict Polymer Physics Just by Reading Synthesis and Processing Prose?

Can large language models predict physical and mechanical polymer properties simply by reading unstructured scientific prose? Polymer performance is rarely determined by chemical structure alone; identical nominal polymers can exhibit drastically different behaviors depending on their synthesis route, processing history, morphology, and testing conditions. Yet, state-of-the-art polymer property models typically rely on structure-only representations -- such as SMILES or molecular graphs -- which strip away this vital experimental context. In this work, we introduce \textbf{PolyLM}, a natural-language-only, process- and condition-aware framework that predicts materials performance directly from full-text literature. By circumventing structural inputs entirely, PolyLM preserves the nuanced, unstructured descriptions of synthesis and processing reported by domain scientists. To train this framework, we curated an unprecedented, literature-scale dataset encompassing 185,000 scientific papers and over 276,400 unique polymer samples across 22 physical, mechanical, and thermal properties. We fine-tuned a massive 9-billion-parameter language model (Qwen3.5-9B) using Low-Rank Adaptation (LoRA) and task-level uncertainty weighting. Evaluated on 68,283 held-out observations, the model achieves remarkably high predictive accuracy, establishing new state-of-the-art benchmarks for complex properties. Across the 22 diverse targets, the model achieves a median R2R^2 of 0.74, with predictions for key thermal, mechanical, and physicochemical properties frequently surpassing an R2R^2 of 0.80. These results unequivocally demonstrate that natural language is a powerful, highly scalable interface for realistic materials performance prediction.
Yuchu Liu, Rui Zhu, Jingwei Xiong +1
Jul 22, 2026cond-mat.mtrl-sci

Generative and multimodal AI for materials prediction and design: Progress, challenges, and perspectives

Artificial intelligence (AI) is accelerating materials prediction and design by enabling efficient exploration of chemical and structural spaces, with particular promise for novel materials discovery. However, novelty in materials discovery encompasses chemical plausibility, structural distinctiveness, property relevance and experimental realisability, making AI-driven novelty claims difficult to substantiate. We introduce a materials property hierarchy, from intrinsic, composition-determined properties to extrinsic, processing-dependent performance, to clarify deployment constraints and distinguish structural, physical and deployment novelty. This framework motivates an evidence-based view of multimodal materials data spanning chemical composition, microstructure, processing, and testing and characterisation, showing that current evidence remains concentrated in composition and idealised structure while heterogeneous, under-represented and weakly integrated modalities limit support for physical and deployment novelty. It also highlights the limitations of benchmarks based mainly on computational labels and proxy novelty criteria. Community-wide standards for data collection, modality alignment and evidence synthesis are needed to support multimodal data construction, process-aware multimodal modelling, feasibility-first generative modelling and deployment-aware benchmarking, so that generative and multimodal AI can design experimentally realisable materials with defensible scientific and practical novelty.
Xianyuan Liu, Charles Anjah, Benjamin E. Jolly +9