Paper Under Review

Agentic AI in Brain Tumor Diagnosis

Nature Computer Vision Medical AI Rejected · Score −3/7
✦
Peer Review Scorecard
Criterion Score Track Rating
Editorial Decision Rejected Score: −3 / 7
1
Summary

This paper proposes an end-to-end pipeline for brain tumor detection and automated clinical report generation by combining YOLOv7 object detection with a Retrieval-Augmented Generation (RAG) large-language-model framework. The system detects and localizes four categories (glioma, meningioma, pituitary tumor, no tumor) from MRI images, maps detection confidence and bounding box attributes to a custom staging heuristic, and generates diagnostic narratives aligned with Indonesian clinical guidelines (PNPK).

YOLOv7 is evaluated on a proprietary MRI dataset of 397 images (319 training / 78 testing) and achieves precision 0.782, recall 0.885, and [email protected] = 0.852. The RAG-LLM component — which retrieves relevant clinical passages and generates structured reports — receives no quantitative evaluation. The paper targets deployment in Indonesian hospital settings where radiologist access is limited.

While the integration concept is timely and the regional clinical localization is a credible contribution, the submission is hampered by methodological flaws in the staging algorithm, a critically small and non-validated dataset, and the complete absence of any evaluation of the generative component.

2
Strengths
Technical Novelty & Innovation
  • Integrating YOLOv7 bounding-box detection with a RAG-LLM clinical reporting engine represents a genuinely novel end-to-end pipeline. Prior work typically addresses detection and reporting as separate systems; coupling them into a single inference flow is architecturally interesting.
  • The use of retrieval-augmented generation to ground clinical narratives in authoritative PNPK guideline passages — rather than unconstrained LLM generation — is a sensible design choice that reflects awareness of hallucination risks in medical AI contexts.
Practical Relevance & Localization
  • Explicitly targeting Indonesian clinical guidelines (PNPK) and low-resource hospital settings distinguishes this work from generic Western-centric benchmarks. The regional localization has genuine applied value.
Detection Performance
  • For the detection sub-task, the reported figures are non-trivial: precision 0.782, recall 0.885, [email protected] = 0.852. These are competitive relative to comparable lightweight detectors on similar four-class MRI benchmarks, though the dataset size limits generalizability claims.
3
Weaknesses
Staging Algorithm: Logical Contradictions
  • The tumor staging heuristic conflates object detection confidence scores with WHO histopathological grading — these are fundamentally different quantities. Detection confidence reflects model certainty about the presence and class of a lesion; WHO grade reflects tissue biology observable only under histology. Equating the two is a clinical validity error that pervades the staging section.
  • The staging rules contain internal contradictions: for example, the boundary conditions across stage categories are either overlapping or leave gaps in the confidence/size space. The algorithm as written would produce inconsistent classifications for a non-trivial proportion of inputs.
Dataset Size & Experimental Validity
  • The evaluation dataset is critically small: 319 training and 78 test images drawn from what appears to be a single source with no patient-level stratification reported. On datasets of this scale, variance in metrics is high and conclusions about generalizability are unsupported.
  • No external validation set or cross-validation is performed. There is no evidence that training and test images are from distinct patient populations, raising the risk of data leakage through patient-level correlations.
  • No ablation study isolates the contribution of individual components. It is unclear whether the staging module, the RAG retrieval, or the LLM generation adds measurable value beyond the YOLOv7 detection baseline.
RAG-LLM Component: Zero Quantitative Evaluation
  • The RAG-LLM pipeline — presented as a core contribution — has no quantitative evaluation whatsoever. There is no factuality assessment, no grounding score (e.g., attribution to retrieved passages), no fluency metric (e.g., BERTScore, ROUGE), no hallucination rate, and no clinical safety evaluation. For a system intended to generate medical diagnostic text, the absence of any evaluation is a critical omission.
Writing & Presentation
  • Terminology is inconsistently applied throughout: "staging" alternates with "grading" and "classification" without precise definitions. Key terms (e.g., "confidence threshold," "bounding box area") that feed the staging algorithm are not formally defined.
  • Related work coverage is thin. Relevant MRI brain tumor detection benchmarks (e.g., BraTS), and recent RAG-for-medical-report works are not cited or compared against.
4
Detailed Comments
Claims Support Negative
The paper's central claim — that the system performs clinically meaningful tumor staging — is not supported by the evidence presented. The staging function operates on YOLOv7 confidence and bounding box area, neither of which is a validated proxy for pathological stage. Furthermore, the 78-image test set is insufficient to demonstrate robust detection performance across the four tumor classes. Claims about real-world deployability in Indonesian hospitals are speculative without a usability study, clinical workflow integration analysis, or prospective pilot data.
Experimental Soundness Negative
The experimental design has three major soundness issues: (1) the test set is too small to yield reliable per-class metrics; the authors report only aggregate mAP, not per-class AP, precision, or recall; (2) there is no patient-level data split — images from the same patient may appear in both training and test sets; (3) the RAG-LLM component is entirely unevaluated, making it impossible to assess whether the generative module is functioning correctly, safely, or at all. A system that is half-evaluated cannot be considered experimentally sound.
Writing Clarity Negative
The manuscript suffers from persistent terminological ambiguity. "Staging" and "grading" are used interchangeably despite referring to distinct clinical processes. The staging algorithm's input variables (confidence threshold, bounding box area, aspect ratio) are introduced in the method section without formal definition or justification of their threshold values. Figure captions are sometimes inconsistent with the corresponding text. The related work section is brief and does not position the paper against the most relevant prior art (BraTS-based detection works, medical RAG literature).
Prior Work Context Neutral
The literature review adequately covers foundational detection models (YOLO family, Faster R-CNN) and basic LLM-for-medicine references. However, it misses several directly relevant threads: (a) BraTS challenge baselines for brain MRI segmentation and detection, (b) recent RAG applications to radiology report generation (e.g., systems grounding reports in structured evidence), and (c) automated diagnostic report generation evaluation frameworks (e.g., CheXpert labeler, RadGraph). The contextual coverage is adequate but not thorough.
Question Importance Positive
The problem being addressed — assisting radiologists in resource-limited settings with automated MRI analysis and report generation — is clinically important and well-motivated. Brain tumor diagnosis is a high-stakes task where delays have significant patient impact. The integration of detection and natural-language reporting is a meaningful architectural direction. The paper is asking the right question; the current execution does not yet meet the standard required to substantiate the answer.
Originality Neutral
The individual components (YOLOv7, RAG, LLM-based medical report generation) are each established. The novelty lies in their combination and the PNPK-aligned localization. This is a reasonable form of applied originality, but the staging heuristic — which could have been a genuinely novel contribution — is undermined by its methodological flaws. If the staging module were rebuilt on clinically valid foundations, the originality of the overall system would be substantially stronger.
Value to Community Negative
In its current form, the paper's value to the research community is limited. The staging algorithm contains errors that would need to be corrected before any replication or extension is possible. The dataset is private and too small to serve as a benchmark. The RAG-LLM component is unevaluated, so its claimed performance cannot be verified or reproduced. A thoroughly revised version that fixes the staging methodology, substantially expands the evaluation, and releases a reproducible evaluation protocol would have meaningful community value.
5
Questions for Authors
  1. How are the staging thresholds (confidence intervals and bounding box area ranges) justified clinically? What is the source for mapping these values to WHO tumor grades, given that WHO grading is defined on histopathological — not imaging-based — criteria?
  2. Were training and test images drawn from distinct patient scans? If multiple images originate from the same patient, was a patient-level split applied to prevent leakage? Please provide dataset provenance and splitting methodology.
  3. Can you provide per-class precision, recall, and AP for each of the four categories (glioma, meningioma, pituitary tumor, no tumor) to allow assessment of per-class performance?
  4. What quantitative evaluation was performed, or can be performed, on the RAG-LLM reports? Specifically: (a) factuality / grounding to retrieved passages, (b) clinical accuracy evaluated by a radiologist, (c) hallucination rate, and (d) comparison against a template-based baseline?
  5. The staging algorithm maps bounding box area to tumor size in cm². What is the pixel-to-cm mapping, and is it consistent across the dataset (i.e., are all MRI scans acquired at the same resolution/FOV)?
  6. Have the PNPK-aligned report templates been reviewed or validated by a board-certified radiologist or neurologist? What process was followed to ensure clinical accuracy of the guideline metadata used in the retrieval index?
  7. Why was external validation on a publicly available brain tumor MRI benchmark (e.g., a subset of BraTS) not included? This would substantially strengthen the generalizability claim.
6
Overall Assessment
Editorial Decision
Rejected — Major Revision Required
The paper addresses a clinically important problem and its PNPK-aligned RAG design shows promise, but the staging algorithm contains clinical validity errors, the evaluation is critically underpowered, and the generative component is entirely unevaluated.
−3
of 7 possible

This paper asks a clinically meaningful question and proposes an architecturally interesting solution. The combination of YOLOv7 detection with RAG-grounded report generation is a timely design pattern, and the regional localization to Indonesian clinical guidelines represents genuine practical value that is underserved in the literature.

However, the manuscript has three critical flaws that preclude publication in its current form. First, the staging algorithm conflates detection model confidence with WHO histopathological grade — a clinical validity error serious enough to make the staging function unreliable as written. Second, the evaluation relies on 78 test images from a single, private dataset with no patient-level splitting, no per-class breakdown, and no comparison to any baseline. Third, the RAG-LLM pipeline — presented as a primary contribution — is entirely unevaluated, rendering its claimed performance unverifiable.

To be considered for publication, the authors should: (1) remove or substantially redesign the staging module with clinically valid foundations; (2) expand evaluation to include a publicly available benchmark, patient-level splits, per-class metrics, and confidence intervals; (3) add quantitative evaluation of the RAG-LLM component covering factuality, fluency, and clinical accuracy reviewed by a domain expert; and (4) clarify all terminological inconsistencies between staging, grading, and classification throughout the manuscript.

A thoroughly revised version addressing these points would represent a meaningful contribution to the medical AI literature and would be worthy of reconsideration.