A Dual-Branch Visual-Textual Model for Contextual Scene Awareness
Integrating Panoptic Segmentation and Scene-Graph Decoding
Project Overview
Conventional captioners name objects but not the relationships between them. This model reasons over structure before it writes a word using panoptic segmentation for pixel-level scene parsing and scene-graph decoding for relational understanding. The dual-branch architecture simultaneously processes visual features and textual semantics for contextually rich captions.
System Architecture
Related Publications (8)
Curated from publications by Professor Dr. Hafiz Ahmad Jalal
RGB-D Scene Classification: A Unified Framework with Vision Transformers and Contextual Models
2024 3rd International Conference on Emerging Trends in Electrical, Control+�n++ +�G�-�, 2024
View on ScholarEnhancing scene understanding using RGB-D visuals and deep learning segmentation models
ETRI, 2026
View on ScholarEnhancing Scene Understanding using RGBD Visuals and Deep Learning Segmentation Models
ETRI, 2026
View on ScholarEnhanced Data Mining and Visualization of Sensory-Graph-Modeled Datasets through Summarization
Sensors, 2024
View on ScholarA Novel Depth Scene Classification with Vision Transformer and ResNet Model
2024 26th International Multi-Topic Conference (INMIC), 1-6, 2024
View on ScholarDynamic Adoptive Gaussian Mixture Model for Multi-Object Detection Over Natural Scenes
ICACS\'24, 2024
View on ScholarScene Understanding and Recognition: Statistical Segmented Model using Geometrical Features and Gaussian Na+�-�ve Bayes
Applied and Engineering Mathematics, 2019
View on ScholarIndividual Sport Activity Detection using Yolo-NAS and Hidden Semi-Markov Model
KHI HTC 2026, 2026
View on Scholar

