MARATTO

review · Journal of risk and financial management

Predicting, Using, and Assessing ESG Signals: A Tripartite Systematic Review of Machine Learning in Sustainable Finance

2026Open accessIbn Tofail University

In plain language

Environmental, Social, and Governance (ESG) ratings increasingly guide capital allocation and corporate strategy, yet they suffer from methodological opacity, divergence across providers, and risks of greenwashing. A systematic review of 127 studies categorises the role of machine learning, deep learning, natural language processing, and explainable artificial intelligence into three areas: predicting scores, using scores, and assessing score credibility. Within the assessment domain, research focuses on reverse-engineering proprietary scoring methods, reconciling conflicting ratings, identifying material industry issues, and detecting greenwashing. The synthesis reveals that commercial ESG ratings consistently assign heavy weight to low-cost aspirational disclosures over verifiable, costly performance metrics, significantly compounding greenwashing risks. Furthermore, high predictive accuracy in reviewed models often stems from non-temporal validation setups rather than robust, transferable out-of-time forecasting capabilities.

Key takeaways

  • Machine learning in sustainable finance is shifting from solely predicting ESG scores to actively evaluating their credibility and construction.
  • Methodological assessment research clusters around reverse-engineering scoring functions, resolving rating divergence, detecting greenwashing, and clustering industry materiality.
  • Commercial ESG ratings often weight low-cost aspirational statements more heavily than verified performance evidence, heightening greenwashing risks.
  • High statistical fit in ESG predictive models frequently reflects weak non-temporal validation rather than genuine forecasting power.

Why it matters

ESG scores directly influence global capital flows, but opaque rating methodologies create substantial risks for investors and society. Showing that existing scores prioritise corporate promises over tangible performance exposes critical vulnerabilities in sustainable finance. Understanding how artificial intelligence can audit and deconstruct these ratings helps market participants, regulators, and civil society challenge misleading corporate sustainability claims.

Commercialisation angle

The reviewed techniques could inform automated compliance software, audit tools, and rating reconciliation platforms for asset managers, financial regulators, and ESG data providers. However, because many underlying models rely on non-temporal validation and reconstruct proprietary inputs rather than generating reliable out-of-time forecasts, these analytical approaches remain at an early-stage research level and require substantial technical validation before direct commercial deployment.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Environmental, Social, and Governance (ESG) ratings increasingly shape capital allocation, corporate strategy, and regulatory oversight, yet their credibility is constrained by methodological opacity, rating divergence, and greenwashing risk. Prior reviews treat machine learning (ML) in ESG as a prediction problem. We identify an emerging research trajectory in which ML is increasingly used not only to consume ESG signals but also to verify their construction and credibility. Drawing on signaling theory, we conduct a PRISMA-guided systematic review of 127 peer-reviewed studies from Scopus and Web of Science to examine how machine learning (ML), deep learning (DL), Natural Language Processing (NLP), and Explainable AI (XAI) are transforming ESG rating analysis. We develop a tripartite framework classifying studies by the functional role of the ESG score: predicted (n = 29), used (n = 57), or assessed (n = 41). Our central contribution is the first synthesis of the methodological-assessment stream, organized into four clusters: XAI reverse-engineering of proprietary scoring functions, divergence reconciliation, greenwashing detection, and unsupervised industry-materiality clustering. The evidence assembled in this stream indicates that ESG ratings weight low-cost aspirational disclosure heavily relative to costly performance evidence, suggesting that greater reliance on aspirational disclosure relative to performance evidence may increase greenwashing risk, consistent with signaling-theory concerns. A study-level validation appraisal further shows that the most extreme fit statistics often arise in target-proximal reconstruction or non-temporal validation settings, cautioning against interpreting high R2 as evidence of transferable out-of-time forecasting.

Research topics

  • Corporate Social Responsibility Reporting
  • Sustainable Finance and Green Bonds
  • Impact of AI and Big Data on Business and Society

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.3390/jrfm19090708

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.