Agentic AI in Brain Tumor Diagnosis
| Criterion | Score Track | Rating |
|---|
This paper proposes an end-to-end pipeline for brain tumor detection and automated clinical report generation by combining YOLOv7 object detection with a Retrieval-Augmented Generation (RAG) large-language-model framework. The system detects and localizes four categories (glioma, meningioma, pituitary tumor, no tumor) from MRI images, maps detection confidence and bounding box attributes to a custom staging heuristic, and generates diagnostic narratives aligned with Indonesian clinical guidelines (PNPK).
YOLOv7 is evaluated on a proprietary MRI dataset of 397 images (319 training / 78 testing) and achieves precision 0.782, recall 0.885, and [email protected] = 0.852. The RAG-LLM component — which retrieves relevant clinical passages and generates structured reports — receives no quantitative evaluation. The paper targets deployment in Indonesian hospital settings where radiologist access is limited.
While the integration concept is timely and the regional clinical localization is a credible contribution, the submission is hampered by methodological flaws in the staging algorithm, a critically small and non-validated dataset, and the complete absence of any evaluation of the generative component.
- Integrating YOLOv7 bounding-box detection with a RAG-LLM clinical reporting engine represents a genuinely novel end-to-end pipeline. Prior work typically addresses detection and reporting as separate systems; coupling them into a single inference flow is architecturally interesting.
- The use of retrieval-augmented generation to ground clinical narratives in authoritative PNPK guideline passages — rather than unconstrained LLM generation — is a sensible design choice that reflects awareness of hallucination risks in medical AI contexts.
- Explicitly targeting Indonesian clinical guidelines (PNPK) and low-resource hospital settings distinguishes this work from generic Western-centric benchmarks. The regional localization has genuine applied value.
- For the detection sub-task, the reported figures are non-trivial: precision 0.782, recall 0.885, [email protected] = 0.852. These are competitive relative to comparable lightweight detectors on similar four-class MRI benchmarks, though the dataset size limits generalizability claims.
- The tumor staging heuristic conflates object detection confidence scores with WHO histopathological grading — these are fundamentally different quantities. Detection confidence reflects model certainty about the presence and class of a lesion; WHO grade reflects tissue biology observable only under histology. Equating the two is a clinical validity error that pervades the staging section.
- The staging rules contain internal contradictions: for example, the boundary conditions across stage categories are either overlapping or leave gaps in the confidence/size space. The algorithm as written would produce inconsistent classifications for a non-trivial proportion of inputs.
- The evaluation dataset is critically small: 319 training and 78 test images drawn from what appears to be a single source with no patient-level stratification reported. On datasets of this scale, variance in metrics is high and conclusions about generalizability are unsupported.
- No external validation set or cross-validation is performed. There is no evidence that training and test images are from distinct patient populations, raising the risk of data leakage through patient-level correlations.
- No ablation study isolates the contribution of individual components. It is unclear whether the staging module, the RAG retrieval, or the LLM generation adds measurable value beyond the YOLOv7 detection baseline.
- The RAG-LLM pipeline — presented as a core contribution — has no quantitative evaluation whatsoever. There is no factuality assessment, no grounding score (e.g., attribution to retrieved passages), no fluency metric (e.g., BERTScore, ROUGE), no hallucination rate, and no clinical safety evaluation. For a system intended to generate medical diagnostic text, the absence of any evaluation is a critical omission.
- Terminology is inconsistently applied throughout: "staging" alternates with "grading" and "classification" without precise definitions. Key terms (e.g., "confidence threshold," "bounding box area") that feed the staging algorithm are not formally defined.
- Related work coverage is thin. Relevant MRI brain tumor detection benchmarks (e.g., BraTS), and recent RAG-for-medical-report works are not cited or compared against.
- How are the staging thresholds (confidence intervals and bounding box area ranges) justified clinically? What is the source for mapping these values to WHO tumor grades, given that WHO grading is defined on histopathological — not imaging-based — criteria?
- Were training and test images drawn from distinct patient scans? If multiple images originate from the same patient, was a patient-level split applied to prevent leakage? Please provide dataset provenance and splitting methodology.
- Can you provide per-class precision, recall, and AP for each of the four categories (glioma, meningioma, pituitary tumor, no tumor) to allow assessment of per-class performance?
- What quantitative evaluation was performed, or can be performed, on the RAG-LLM reports? Specifically: (a) factuality / grounding to retrieved passages, (b) clinical accuracy evaluated by a radiologist, (c) hallucination rate, and (d) comparison against a template-based baseline?
- The staging algorithm maps bounding box area to tumor size in cm². What is the pixel-to-cm mapping, and is it consistent across the dataset (i.e., are all MRI scans acquired at the same resolution/FOV)?
- Have the PNPK-aligned report templates been reviewed or validated by a board-certified radiologist or neurologist? What process was followed to ensure clinical accuracy of the guideline metadata used in the retrieval index?
- Why was external validation on a publicly available brain tumor MRI benchmark (e.g., a subset of BraTS) not included? This would substantially strengthen the generalizability claim.
This paper asks a clinically meaningful question and proposes an architecturally interesting solution. The combination of YOLOv7 detection with RAG-grounded report generation is a timely design pattern, and the regional localization to Indonesian clinical guidelines represents genuine practical value that is underserved in the literature.
However, the manuscript has three critical flaws that preclude publication in its current form. First, the staging algorithm conflates detection model confidence with WHO histopathological grade — a clinical validity error serious enough to make the staging function unreliable as written. Second, the evaluation relies on 78 test images from a single, private dataset with no patient-level splitting, no per-class breakdown, and no comparison to any baseline. Third, the RAG-LLM pipeline — presented as a primary contribution — is entirely unevaluated, rendering its claimed performance unverifiable.
To be considered for publication, the authors should: (1) remove or substantially redesign the staging module with clinically valid foundations; (2) expand evaluation to include a publicly available benchmark, patient-level splits, per-class metrics, and confidence intervals; (3) add quantitative evaluation of the RAG-LLM component covering factuality, fluency, and clinical accuracy reviewed by a domain expert; and (4) clarify all terminological inconsistencies between staging, grading, and classification throughout the manuscript.
A thoroughly revised version addressing these points would represent a meaningful contribution to the medical AI literature and would be worthy of reconsideration.