Transforming Oncology Imaging with Foundation Models for Intelligent Radiology Report Generation
DOI:
https://doi.org/10.65477/ijrems.v2.i1.01Keywords:
Foundation models; Automated radiology report generation; Oncology imaging; Large language models; Vision-language models; Artificial intelligence; Medical imaging; Precision oncology.Abstract
Radiology reports are the primary means of communication between radiologists and referring clinicians, providing essential information for disease diagnosis, treatment planning, therapeutic response assessment, and longitudinal patient management. In oncology, the increasing volume, complexity, and multimodal nature of imaging examinations have created a growing demand for reporting systems that are both accurate and efficient. Recent advances in artificial intelligence (AI), particularly the emergence of foundation models, have significantly transformed automated radiology report generation. Trained on large-scale multimodal datasets, foundation models learn generalized representations that can be adapted to a wide range of clinical tasks, enabling the generation of coherent, context-aware, and clinically meaningful radiology reports through the integration of medical imaging and natural language processing.
This review explores the evolution of automated radiology report generation, tracing its progression from traditional rule-based systems and convolutional neural network (CNN)-based approaches to transformer architectures, vision-language models, and multimodal large language models (LLMs). Particular emphasis is placed on their applications in oncology imaging, including computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), and hybrid imaging modalities. The review further examines major foundation model architectures, publicly available datasets, evaluation benchmarks, clinical applications, and the advantages and limitations of current methodologies. In addition, key challenges associated with clinical implementation—including model hallucination, limited explainability, data heterogeneity, privacy protection, regulatory compliance, and the need for rigorous prospective validation—are critically discussed.
Current evidence indicates that foundation models outperform earlier AI-based approaches by improving report quality, contextual understanding, linguistic coherence, and cross-domain generalizability. Nevertheless, several technical and clinical challenges must be addressed before these systems can be safely integrated into routine oncology practice. Future research should prioritize the development of domain-specific multimodal foundation models, federated learning frameworks for privacy-preserving collaboration, explainable AI techniques to enhance clinical trust, and large-scale prospective validation studies. As these technologies continue to mature, foundation models have the potential to transform radiology workflows by improving reporting efficiency, consistency, diagnostic accuracy, and clinical decision support, ultimately contributing to more precise and personalized cancer care.

