AN NOVEL DEEP LEARNING FRAMEWORK INTEGRATING VISIONTRANSFORMERS AND MULTI-SCALE DENSE CAPSNETS FOR EARLY-STAGEHEPATOCELLULAR CARCINOMA DETECTION
Dharmarajan, K and Abirami, K and Haripriya, T (2026) AN NOVEL DEEP LEARNING FRAMEWORK INTEGRATING VISIONTRANSFORMERS AND MULTI-SCALE DENSE CAPSNETS FOR EARLY-STAGEHEPATOCELLULAR CARCINOMA DETECTION. International Journal of Engineering Technology Research & Management (IJETRM), 10 (6). pp. 243-248. ISSN 2456-9348
Jun-2026-12-1781289048-AN-JUNE2026-23.pdf
Download (225kB)
Abstract
Primary liver cancer, predominantly Hepatocellular Carcinoma (HCC), represents one of the leading causes ofoncological mortality globally. Early detection via Computed Tomography (CT) and Magnetic ResonanceImaging (MRI) is paramount for improving patient survival rates, yet manual tumor segmentation andclassification remain highly subjective and error-prone due to variations in tumor morphology, fuzzyboundaries, and low-contrast clinical scans. To address these critical limitations, this paper proposes aninnovative hybrid deep learning framework, named HepatoX-Net. The framework combines the globalcontextual representation capabilities of Vision Transformers (ViTs) with the spatial-hierarchical andorientation-preserving strengths of Multi-Scale Dense Capsule Networks (Dense CapsNets). The architectureoperates in a dual-pathway manner: Pathway A employs a customized Swin Transformer backbone with shiftedwindowing mechanisms to capture long-range non-local dependencies and textural irregularities across 3Dmultiphasic imaging volumes, while Pathway B utilizes a novel dense routing Capsule Network to extractprecise spatial configurations, geometric transformations, and boundaries of focal liver lesions. A DynamicCross-Attention Fusion (DCAF) module bridges the two pathways, adaptively weighting features to optimizesemantic representation while discarding noisy artifacts. Extensive experimental evaluations were executed onthe public LiTS2017 (Liver Tumor Segmentation Challenge) dataset and an institutional multi-phasic CT/MRIclinical repository containing 1,420 validated cases collected across 2025 and early 2026. The proposedHepatoX-Net framework achieved a state-of-the-art classification accuracy of 98.42%, a sensitivity of 97.89%,and a Dice Similarity Coefficient (DSC) of 0.941 for complex multi-class liver lesion segmentation. Our modelsubstantially outpaced baseline architectures including U-Net++, ResNet-151, and standard VisionTransformers, demonstrating robust generalization, high structural reliability, and clinical viability as acomputer-aided diagnostic system.
| Item Type: | Article |
|---|---|
| Subjects: | Computer Applications > Artificial Intelligence |
| Domains: | Computer Science |
| Depositing User: | Mr IR Admin |
| Date Deposited: | 06 Aug 2026 09:52 |
| Last Modified: | 01 Sep 2026 11:26 |
| URI: | https://ir.vistas.ac.in/id/eprint/21994 |
