
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
A REVIEW OF DEVELOPMENT OF A HIERARCHICAL ATTENTION-GUIDED DEEP CONVOLUTIONAL NETWORK FOR CONTEXT-AWARE IMAGE UNDERSTANDING WITH DYNAMIC FEATURE SUPPRESSION IN PYTHON
Km. Mahima Verma1, Mrs. Arifa Khan2
1Master of Technology, Computer Science and Engineering, Lucknow Institute of Technology, Lucknow, India 2Assistant Professor, Department of Computer Science and Engineering, Lucknow Institute of Technology, Lucknow, India ***
Abstract - Therapidevolutionofdeepconvolutionalneural networks (CNNs) has significantly advanced image understanding.However,traditionalmodelsoftenstruggleto capture long-range dependencies and salient contextual information, treating all spatial and channel-wise features uniformly. This uniform processing leads to computational inefficiency and suboptimal performance in complex scenes wherecontextiskey.Theemergenceofhierarchicalattention mechanisms and dynamic feature suppression offers a promising paradigm to address these limitations. This paper presentsasystematicreviewofattention-guideddeeplearning architectures designed for context-aware image understanding. We survey the landscape from soft and hard attention to modern Transformers and dynamic gating mechanisms. We synthesize the literature into a taxonomy, discuss the theoretical underpinnings of feature suppression, and analyze the integration of these components into hierarchical networks. We identify a critical research gap in thejointoptimizationofattentionforboth selection (whatto look at) and suppression (what to ignore). The review concludes by outlining open challenges and proposing future research directions, including the development of unified frameworks and efficient Python-based implementations for real-world applications.
Key Words: Attention Mechanisms; Context-Aware Image Understanding; Deep Learning; Dynamic Feature Suppression; Hierarchical Neural Networks; Computer Vision
1. INTRODUCTION
1.1. Background and Motivation
The field of computer vision has witnessed a remarkable evolution over the past decade, transitioning from basic imageclassificationtaskstocomplexsceneunderstanding problemssuchassemanticsegmentation,imagecaptioning, and visual question answering. This journey began with groundbreaking work on large-scale image classification usingdeepconvolutionalneuralnetworks(Krizhevskyetal., 2012), which demonstrated that hierarchical feature learningcouldachieveunprecedentedaccuracyondatasets like ImageNet. Subsequent advances in network architecture,includingtheintroductionofVGG(Simonyan
andZisserman,2015)andresiduallearninginResNet(Heet al.,2016),progressivelypushedtheboundariesofwhatdeep modelscouldaccomplish.Thesefoundationaldevelopments paved the way for more sophisticated vision tasks that require not merely object recognition but a holistic understanding of visual scenes, including relationships between objects, spatial configurations, and semantic context(Longetal.,2015;Renetal.,2017).
Despite these advances, standard convolutional neural networkspossessinherentlimitationsthatconstraintheir ability to achieve genuine context-aware understanding. Traditional CNNs operate with fixed receptive fields determined by kernel sizes and network depth, which fundamentally limits their capacity to capture long-range dependenciesandglobal contextual information(Wang et al.,2018).Whilestackingmultipleconvolutionallayerscan theoretically expand the receptive field, in practice, the effective receptive field is often much smaller than the theoreticalmaximumduetotheconcentrationofinfluence in central regions (Luo et al., 2016). Furthermore, conventionalCNNsprocessallspatiallocationsandfeature channelsuniformly,treatingeveryregionofanimagewith equalimportanceregardlessofitsrelevancetothetaskat hand. This uniform processing paradigm leads to two significant shortcomings: computational inefficiency, as resources are expended on processing irrelevant background regions, and suboptimal performance in complexsceneswherecontextualrelationshipsarecrucial foraccurateinterpretation(Huetal.,2018).
The "context-awareness" problem in computer vision fundamentally concerns the challenge of understanding visual elements not in isolation but in relation to their surroundings. Recognising an object often requires understanding the scene in which it appears a small, cylindricalobjectmightbeidentifiedasacupwhensituated onadiningtablebutinterpreteddifferentlywhenfoundina bathroom setting (Oliva and Torralba, 2007). Similarly, interpreting human actions requires understanding the objects involved and the environment where the action occurs.Thiscontextualreasoning,whichcomesnaturallyto human perception, proves remarkably challenging for artificialvisionsystems.Earlyapproachestoincorporating context involved multi-scale architectures and spatial

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
pyramid pooling (Lazebnik et al., 2006; He et al., 2015), whichaggregatedinformationatmultipleresolutions,and atrousconvolutiontechniquesthatexpandedreceptivefields without increasing parameters (Chen et al., 2017). While thesemethodsrepresentedimportantstepstowardcontextaware processing, they lacked the adaptive, selective mechanismscharacteristicofhumanvisualattention.
1.1.1. The Evolution from Classification to Scene Understanding
Theprogressionfromimageclassificationtocomprehensive scene understanding marks a fundamental shift in the objectivesofcomputervisionsystems.Imageclassification tasksrequireassigninga singlelabel toan entireimage,a problem that can often be solved by identifying the most salientobjectpresent(Russakovskyetal.,2015).Incontrast, sceneunderstandingdemandsdetailedinterpretationofall elements within an image, their relationships, and the overall semantic context. Semantic segmentation requires pixel-levelclassification,whereeachpixelmustbeassigned to a category, necessitating both local detail preservation andglobalcontextualcoherence(Longetal.,2015).Image captioninggoesfurther,requiringthegenerationofnatural language descriptions that capture not only the objects present but also their attributes, actions, and interactions (Vinyalsetal.,2015).Visualquestionansweringchallenges modelstoreasonaboutimagesbasedonnaturallanguage questions, often requiring sophisticated reasoning about spatial relationships, counting, and commonsense knowledge(Antoletal.,2015).Eachofthesetasksdemands forms of contextual understanding that exceed the capabilitiesofstandardCNNsoperatingwithfixedreceptive fields and uniform feature processing, motivating the development of more sophisticated attention-guided architectures.
1.2.The Rise of Attention and Dynamic Mechanisms
The concept of attention in neural networks draws direct inspiration from the human cognitive system, which possessesaremarkableabilitytofocusprocessingresources on salient regions of the visual field while suppressing irrelevant information (Itti et al., 1998). In human perception, attention operates through a combination of bottom-up,saliency-drivenmechanismsandtop-down,taskguided selection, enabling efficient processing of complex visual scenes despite limited neural resources. This biologicalparadigmhasprovenhighlyinfluentialinartificial intelligence, leading to the development of computational attention mechanisms that enable deep networks to selectively focus on informative features while ignoring distractors(Hassaninetal.,2024).
The history of attention mechanisms in neural networks traces an interesting trajectory from vision to natural languageprocessingandbacktovision.Earlycomputational modelsofvisualattention,suchasthesaliency-basedsystem
proposed by Itti et al. (1998), demonstrated that biologically-inspired attention could efficiently identify importantregionsincomplexscenes.Thisworkestablished foundational principles of computational attention that wouldlaterinfluencedeeplearningapproaches.Asignificant milestoneoccurredwhenMnihetal.(2014)introducedthe recurrentattentionmodel(RAM),whichfirstincorporated attention mechanisms into recurrent neural networks for imageclassification.Thisworkdemonstratedthatattention could enable networks to process images sequentially, focusingondifferentregionsovertime,muchlikehumaneye movements.
The transformative impact of attention in deep learning, however,ismostfrequentlyassociatedwithitsapplication in natural language processing. Bahdanau et al. (2015) introduced attention mechanisms to neural machine translation, enablingthemodel toaligntarget wordswith relevant source words during translation. This work addressedthefundamentallimitationoffixed-lengthcontext vectorsinencoder-decoderarchitecturesanddemonstrated thatattentioncouldsignificantlyimproveperformance on sequence-to-sequencetasks.Thesubsequentintroductionof theTransformerarchitecture(Vaswanietal.,2017),which dispensed with recurrent and convolutional processing entirely in favour of pure self-attention, revolutionised naturallanguageprocessingandestablishedattentionasa fundamentalbuildingblockofmoderndeeplearning.
ThesuccessofattentioninNLPcatalyseditsreintroduction tocomputervisionwithrenewedvigour.Thesqueeze-andexcitation network (SENet) (Hu et al., 2018) introduced channel attention, enabling networks to adaptively recalibrate channel-wise feature responses by modelling interdependencies between channels. This simple yet effective mechanism could be incorporated into any CNN architecture and yielded consistent performance improvements across multiple tasks. The convolutional blockattentionmodule(CBAM)(Wooetal.,2018)extended this idea by sequentially applying channel and spatial attention, demonstrating that complementary attention mechanisms could work synergistically. Non-local neural networks(Wangetal.,2018)broughtself-attentiontovision by enabling each position to attend to all other positions, capturing long-range dependencies that traditional convolutionscouldnot.Finally,theVisionTransformer(ViT) (Dosovitskiy et al., 2021) applied pure transformer architectures to image classification, demonstrating that with sufficient data, attention-based models could outperform convolutional networks on vision tasks. This trajectory of development has established attention as an indispensable component of modern computer vision architectures.
1.2.1. Defining Dynamic Feature Suppression
Dynamic feature suppression represents a distinct yet complementaryconcepttoattention,concernedspecifically

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
with the active inhibition of irrelevant or noisy features rather than merely weighting them less heavily. While attentionmechanismstypicallycomputeimportanceweights thatscalefeatureresponses,suppressionmechanismscan entirely eliminate or gate the flow of information from unimportantchannelsorspatiallocations(Gaoetal.,2019). This distinction carries important implications for both model performance and computational efficiency. The importance of active suppression arises from the observation that the saliency of neurons in convolutional layers is highly input-dependent a channel that proves criticalforoneimagemaybenearlyirrelevantforanother (Gaoetal.,2019).Staticpruningmethodsthatpermanently remove channels based on average importance across a datasetinevitablysacrificethecapabilitytohandleinputsfor which those pruned channels would have been essential. Dynamicsuppressionpreservesthefullnetworkstructure while skipping computation for unimportant channels at run-time, maintaining model capacity while reducing computationalcost.
1.3. Scope and Contributions of this Review
Thispaperpresentsasystematicreviewofattention-guided deep learning architectures for context-aware image understanding, with a specific focus on hierarchical networksanddynamicfeaturesuppressionmechanisms.Itis important to state explicitly that this work constitutes a review paper rather than a novel methodological contribution.Ourobjectiveistosynthesisetheextensiveand rapidly growing literature on attention mechanisms in computer vision, identify patterns and principles that transcendindividualcontributions,andprovideastructured analysis that can guide future research in this area. The reviewadoptsacriticalperspective,evaluatingnotonlythe achievements of existing work but also identifying limitations, open challenges, and promising directions for furtherinvestigation.
The scope of this review is carefully delimited to ensure focused and coherent analysis. We concentrate on hierarchical,multi-scaleattentionmechanismsthatoperate atmultiplelevelsoffeatureabstraction,fromlow-leveledges and textures to high-level semantic concepts. This hierarchical perspective is essential for context-aware understanding, as contextual information manifests at multiple scales local context informs object boundaries, whileglobalcontextdeterminesscenesemantics(Chenetal., 2018). We specifically examine attention-guided deep networksdesignedforcontext-awareimageunderstanding, including applications in semantic segmentation, scene parsing,imagecaptioning,andvisualquestionanswering.A distinctivefocusofthisreviewistheexaminationofdynamic featuresuppressionmechanisms,whichhavereceivedless systematic attention in previous surveys despite their growingimportanceforbothefficiencyandrobustness.We considersuppressionmechanismsincludingchannelgating,
spatial pruning, and dynamic convolution, analysing their relationship to attention and their role in comprehensive context-awaresystems.
1.3.1. Primary Contributions of This Review
This review makes four primary contributions to the literatureonattentionmechanismsincomputervision.First, we develop a novel taxonomy of attention-suppression mechanisms that organises the literature according to fundamental design dimensions rather than superficial architectural differences. This taxonomy distinguishes betweensoftandhardattention,channel,spatial,andmixed attention,andfurthercategorisessuppressionmechanisms accordingtotheirgranularity(channel-wise,spatial-wise,or element-wise) and their dynamic or static nature. By providing a structured framework for understanding the relationshipsbetweendifferentapproaches,thistaxonomy aimstofacilitatecomparisonandguidearchitecturaldesign decisions.
Second, we provide a synthesis of architectural design patterns that recur across successful attention-guided networks.Ratherthanpresentinganexhaustivecatalogueof individualmodels,weidentifycommonbuildingblocksand designprinciples suchastheencoder-decoderparadigm with attention in skip connections, the integration of selfattention for long-range dependency modelling, and the combinationofcomplementaryattentionmechanisms.This synthesisaimstodistilpracticalinsightsthatcaninformthe development of new architectures for context-aware understanding.
2. LITERATURE REVIEW
2.1. Foundational Concepts: Context in Computer Vision
The concept of context has long been recognised as fundamental to visual perception, both in biological and artificialvisionsystems.Incomputervision,contextrefersto theinformationsurroundingatargetelementthataidsinits interpretation, encompassing everything from local pixel neighbourhoods to global scene semantics (Oliva and Torralba,2007).Theintegrationofcontextualinformation addresses a fundamental limitation of isolated object recognition:manyobjectsareinherentlyambiguouswhen viewed in isolation and only become identifiable through their relationship with the surrounding environment. For instance,asmallcylindricalobjectmightbeinterpretedasa cup when situated on a dining table, but as a toothbrush holder when found in a bathroom setting. This contextual disambiguationisessentialforrobustimageunderstanding systems that must operate in unconstrained real-world environments.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
2.1.1. Global versus Local Context
Early approaches to incorporating context in deep neural networks focused on multi-scale architectures that could simultaneouslycapturebothfine-grainedlocaldetailsand coarse global scene structure. The fundamental insight underlying these approaches is that different semantic information manifests at different spatial scales: object boundaries require high-resolution local context, while scenecategoryandoverallspatiallayoutarebestcapturedat lowerresolutionswithlargerreceptivefields(Farabetetal., 2013).
Spatial pyramid pooling represented one of the earliest systematic attempts to address multi-scale context aggregation.Lazebniketal.(2006)introducedtheconceptof pyramid matching in the context of scene recognition, partitioning images into increasingly fine subregions and computing histograms of local features within each subregion. This approach was later adapted for deep learningbyHeetal.(2015),whoproposedspatialpyramid poolinginconvolutionalneuralnetworkstogeneratefixedlengthrepresentationsregardlessofinputsize,enablingthe network to aggregate information at multiple scales simultaneously.
Aparallellineofdevelopmentemergedfromtherecognition that standard convolutional operations with fixed kernels could not efficiently capture context at multiple scales without incurring prohibitive computational costs. Atrous convolution,alsoknownasdilatedconvolution,providedan elegant solution to this problem by introducing a dilation rate parameter that controls the spacing between kernel weights(Chenetal.,2018).Thismechanismallowsfiltersto operate with enlarged receptive fields without increasing parameter count or computational complexity, effectively enablingthenetworktocapturemulti-scalecontextthrough parallelbrancheswithdifferentdilationrates.AsChenetal. (2018) demonstrated, atrous convolution proves particularly effective for dense prediction tasks where maintaining spatial resolution while expanding receptive fieldsiscritical.
2.1.2. Context via Recurrent Models
While multi-scale convolutional approaches effectively capturespatialcontextatmultipleresolutions,theyprocess all image regions simultaneously and uniformly. An alternativeparadigmdrawsinspirationfromhumanvisual perception, which sequentially samples the visual environment through saccadic eye movements, focusing attentiononsalientregionswhilemaintainingaglobalscene representation.Recurrentneuralnetworks(RNNs)andtheir variants, particularly Long Short-Term Memory (LSTM) networks,providedanaturalframeworkforimplementing suchsequentialprocessingindeeplearningsystems.
Therecurrentattentionmodel(RAM)introducedbyMnihet al. (2014) represented a seminal contribution to this direction.RAMprocessesimagesbysequentiallyselecting and processing small image regions, building up a representationovertimethrougharecurrentnetworkthat learnsbothwheretolookandhowtointerpretthegathered information. This approach demonstrated that sequential attention could achieve competitive classification performancewhileprocessingonlyafractionoftheimage pixels,highlightingthecomputationalefficiencybenefitsof selectivecontextualsampling.
2.2. Attention Mechanisms in Vision: A Taxonomy
The success of attention mechanisms in neural machine translation (Bahdanau et al., 2015) catalysed extensive researchintoattention-basedmodelsforcomputervision, yieldingadiverselandscapeofmechanismsthatoperateat different levels of representation and employ various computational strategies. Organising this literature into a coherent taxonomy requires considering multiple dimensions:thedomainofapplication(spatial,channel,or hybrid), the nature of attention computation (soft deterministicorhardstochastic),andthearchitecturallevel at which attention operates (single-scale or hierarchical). This section develops such a taxonomy to provide a structuredframeworkforunderstandingtherelationships betweendifferentattentionapproaches.
2.2.1. Soft versus Hard Attention
Thedistinctionbetweensoftandhardattentioncentreson the differentiability and stochasticity of the attention mechanism, with profound implications for training methodology and architectural design. Soft attention mechanisms computea weightedcombinationofall input elements, with weights typically derived from a differentiable compatibility function, enabling end-to-end training via standard backpropagation. Hard attention mechanisms,incontrast,selectasubsetofinputelements throughstochasticsampling,renderingtheselectionprocess non-differentiable and necessitating alternative training approachessuchasreinforcementlearning.
The Selective Kernel Network (SKNet) (Li et al., 2019) extendedchannelattentionbyintroducingdynamickernel selection based on attention mechanisms. SKNet employs multiple branches with different kernel sizes, aggregates information across branches through element-wise summation, and then applies attention mechanisms to selectivelyemphasisefeaturesfromdifferentkernelscales based on input content. This enables the network to adaptively adjust its receptive field size depending on the scaleofobjectspresentintheinput,effectivelylearningto select appropriate convolutional kernel sizes through soft attention.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
2.2.2. Hierarchical and Multi-scale Attention
While single-scale attention mechanisms can effectively highlight important features at a fixed resolution, natural images contain structures at multiple scales that require attentionoperatingatcorrespondinglevelsofabstraction. Hierarchicalattentionmechanismsaddressthisrequirement by applying attention at different levels of feature hierarchies,fromlow-leveledgeandtexturerepresentations in early network layers to high-level semantic representationsindeeperlayers.
PyramidAttentionNetworks(PAN)introducedbyLietal. (2018) explicitly address multi-scale attention through a pyramidal structure that applies attention at multiple feature resolutions. PAN first generates attention maps at multiplescales,thenusesthesemapstoweightfeaturemaps atcorrespondingscales,andfinallyfusestheweightedmultiscale features through upsampling and summation. This design enables the network to capture both fine-grained details through high-resolution attention and broad contextual relationshipsthroughlow-resolutionattention, with information flowing both bottom-up and top-down throughtheattentionpyramid.
2.2.3.
Self-Attention and Transformers
TheintroductionoftheTransformerarchitecture(Vaswani etal.,2017)innaturallanguageprocessing revolutionised sequence modelling by demonstrating that pure selfattention,withoutrecurrentorconvolutionalcomponents, could achieve state-of-the-art performance while offering superiorparallelisationandtheabilitytocapturelong-range dependencies.Thissuccesscatalysedextensiveresearchinto adapting Transformer architectures for computer vision, fundamentally reshaping the landscape of visual representationlearning.
2.3. Dynamic Feature Suppression: Gating and Selection
Whileattentionmechanismsfocusonamplifyingimportant features, an equally important aspect of selective informationprocessingisthesuppressionofirrelevant or distracting information. Dynamic feature suppression encompasses mechanisms that actively gate or prune features based on input content, preventing irrelevant information from propagating through the network and potentiallycorruptingrepresentations.Thissectionreviews approaches to dynamic suppression at different levels of representation, drawing inspiration from both computational efficiency considerations and biological principlesofselectiveattention.
Thebiologicalinspirationfordynamicsuppressionderives fromneurocognitiveresearchdemonstratingthatthebrain actively suppresses distracting information rather than merelyfailingtoattendtoit.DiBelloetal.(2021)providea
unified neurocognitive framework describing how the prefrontalcorteximplementsbothproactivesuppression strategically prioritising task-relevant information at the expense of irrelevant information and reactive suppression interrupting the ongoing processing of distractors that have captured attention. This dualmechanism perspective highlights that effective selective processing requires both preventing distraction before it occursandterminatingdistractionwhenitbreaksthrough initialfilters.
2.3.1. Channel-wise Suppression
Channel-wise suppression mechanisms operate along the channeldimensionofconvolutionalfeaturemaps,selectively inhibitingorgatingentirefeaturechannelsbasedontheir estimatedrelevancetothecurrentinput.TheSqueeze-andExcitation Network (SENet) (Hu et al., 2018) can be interpreted as implementing a form of soft channel suppression: the excitation operation produces channel weightsthatcanapproachzeroforunimportantchannels, effectively suppressing their contribution to subsequent layers.However,SENet'ssuppressionismultiplicativeand continuous channels with low weights contribute minimally but are not completely eliminated, and the computation required to process these channels remains unchanged.
More aggressive channel-wise suppression mechanisms explicitly zero out or skip computation for unimportant channels, achieving both representational benefits and computational savings. Gao et al. (2019) introduced GaterNet, a dual-path architecture where a lightweight "gater" network learns to generate binary gates that selectively activate or deactivate channels in a main "backbone"network.Thegaternetworkprocessesthesame inputasthebackbonebutwithreducedcapacity,learningto predictwhichchannelswillbeusefulforthecurrentinput. Channels with zero gates are completely skipped during backbone computation, reducing both the number of operationsandthememoryfootprint.Thisdynamicchannel gatingenablesthenetworktoadaptitseffectivecapacityto each input, allocating more computation to challenging exampleswhileprocessingsimpleexamplesefficiently.
2.3.2. Spatial-wise Suppression
Spatial-wisesuppressionmechanismsoperateonthespatial dimensionsoffeaturemaps,selectivelyinhibitingorpruning responses at specific spatial locations. This form of suppressionrecognisesthatinnaturalimages,manyspatial regionscontainbackgroundorirrelevantcontentthatcanbe safelyignoredwithoutcompromisingtaskperformance.By suppressingsuchregions,spatialattentionmechanismscan focuscomputationalresourcesoninformativeareaswhile reducingtheinfluenceofdistractingbackgroundpatterns.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Theconnectionbetweenspatialsuppressionandsparsityis fundamental: suppressing spatial locations effectively introducesspatialsparsityinfeaturerepresentations,with only a subset of locations propagating significant information to subsequent layers. This sparsity can be exploited for computational efficiency through sparse convolution operations that only process active locations (Graham et al., 2018), though practical implementations oftenrequirespecialisedhardwareorsoftwaresupportfor irregularcomputationpatterns.
2.3.3. Dynamic Convolution
Dynamicconvolutionrepresentsamorefundamentalform of feature adaptation than attention or gating, where the convolutional kernels themselves are generated or assembleddynamicallybasedoninputcontent.Ratherthan selectingwhichfeaturestoemphasiseorsuppresswithina fixedkernelrepresentation,dynamicconvolutionadaptsthe kernelsto the input, enablingcontent-parametric filtering thatcancaptureinput-specificpatterns.
CondConv (Yang et al., 2019) introduces conditionally parameterised convolutions where kernel weights are computedasalinearcombinationofmultipleexpertkernels, with combination weights predicted from the input. Specifically, for each input example, CondConv computes routingweightsthroughalightweightgatingfunction,then combines multiple convolutional kernels by weighted summation to produce an input-specific kernel. This approach increases model capacity without increasing inference cost proportionally, as the kernel combination requiresonlyO(K)O(K)operationswhereKKisthenumber of experts, while the subsequent convolution uses the combined kernel of standard size. Yang et al. (2019) demonstrate that CondConv improves accuracy across multiplearchitectureswhilemaintainingefficientinference, effectively learning to adapt convolutional filters to input content.
2.4. Integration into Downstream Tasks: ContextAware Understanding
Theultimatetestofattentionandsuppressionmechanisms liesintheirabilitytoimproveperformanceondownstream tasksrequiringgenuinecontext-awareunderstanding.This section reviewshowthetechniquesdiscussedabovehave been integrated into architectures for semantic segmentation, image captioning, scene parsing, and visual question answering, examining both the specific design choices that enable effective context modelling and the empiricalbenefitsdemonstratedonstandardbenchmarks.
2.4.1. Semantic Segmentation
Semanticsegmentation,thetaskofassigningaclasslabelto everypixelinanimage,hasservedasaprimarytestbedfor context-awarearchitecturesduetoitsinherentrequirement
forintegratinglocaldetailwithglobalsceneunderstanding. Accurate segmentation demands both precise boundary localisation, requiring high-resolution features, and consistent labelling across object instances, requiring contextualreasoningaboutobjectrelationshipsandscene semantics.
OCRNet(Object-ContextualRepresentations)introducedby Yuanetal.(2020)representsasignificantadvanceincontext modelling for segmentation by explicitly representing object-contextual information. The key insight underlying OCRNetisthatthemostrelevantcontextforapixelcomes from other pixels belonging to the same object category, ratherthanfromallpixelsindiscriminately.Toexploitthis, OCRNetfirstestimatesacoarseobjectregionrepresentation throughasegmentationhead,thencomputespixel-to-region relationships that weight the contribution of each object regiontoeachpixel'srepresentation.Thisapproachenables each pixel to attend selectively to regions containing the same object category, effectively implementing a form of object-aware contextual aggregation. Yuan et al. (2020) demonstrate that OCRNet achieves state-of-the-art performance on multiple segmentation benchmarks, confirmingthevalueofobject-awarecontextmodelling.
2.4.2. Scene Parsing and Image Captioning
Scene parsing, which involves segmenting an image into semantically meaningful regions corresponding to objects and their parts, demands even richer contextual understanding than semantic segmentation alone. The Pyramid Scene Parsing Network (PSPNet) introduced by Zhaoetal.(2017)addressesthisthroughpyramidpooling modulesthat aggregate context atmultiplescales.PSPNet appliespoolingoperationsatdifferentgridscales,producing pooled representations that capture global, regional, and local context, then upsamples and concatenates these representationstoformacomprehensivecontextualfeature. Thisdesignenablesthenetworktoincorporateinformation fromtheentirescenewhenparsingeachpixel,reducingthe likelihoodoflocalconfusion(e.g.,mistakingaboatforacar whenbothappearnearwater).
Image captioning requires generating natural language descriptions that capture not only the objects present but also their attributes, relationships, and activities. The UpDownattentionmodel(Andersonetal.,2018)introduceda two-levelattentionmechanismthatfirstattendstoobjects (throughbottom-upregionproposals)andthenattendsto relationshipsbetweenobjects(throughtop-downattention conditionedonthecurrentlanguagestate).Thisapproach recognises that effective caption generation requires both identifying salient objects and understanding how they relatetoeachother,withattentionoperatingatbothlevels. Anderson et al. (2018) demonstrated state-of-the-art performance on multiple captioning benchmarks, establishingthevalueofhierarchicalattentionthatcombines bottom-upsaliencywithtop-downtaskguidance.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
2.4.3. Visual Question Answering
VisualQuestionAnswering(VQA)presentsperhapsthemost demandingtestofcontext-awareunderstanding,requiring systemstoanswernaturallanguagequestionsaboutimages by integrating visual perception, language understanding, and commonsense reasoning. Questions can range from simple object identification ("What colour is the car?") to complex relational reasoning ("How many children are standing to the left of the woman holding an umbrella?"), each requiring different forms of visual attention and contextualaggregation.
StackedAttentionNetworks(SAN)introducedbyYangetal. (2016)addressVQAthroughmultipleattentionlayersthat progressivelyrefinethefocusonimageregionsrelevantto the question. The first attention layer identifies coarse regions potentially containing relevant information, while subsequent layers refine this focus based on the question context and previously attended regions. This multi-step attentionprocessmirrorshumanvisual searchbehaviour, whereinitialglancesidentifypromisingregionsfollowedby closer inspection of detailed content. Yang et al. (2016) demonstratedthatstackedattentionsignificantlyimproves VQAaccuracy,particularlyforquestionsrequiringmulti-step reasoning.
3.ANALYSISOF ARCHITECTURALDESIGNPATTERNS
3.1. The Encoder-Decoder Paradigm
The encoder-decoder architecture has emerged as a dominantparadigmfordensepredictiontasksincomputer vision,includingsemanticsegmentation,depthestimation, andimage-to-imagetranslation.Thisdesignpatternconsists of an encoder network that progressively reduces spatial resolutionwhileincreasingfeaturedimensionality,capturing increasingly abstract representations, and a decoder network that progressively recovers spatial resolution to produce output at the original input dimensions (Badrinarayanan et al., 2017). The encoder-decoder structurenaturallysupportsmulti-scalefeaturelearningand has proven particularly effective when combined with attention mechanisms that selectively emphasise relevant informationatdifferentstagesoftheprocessingpipeline.
The integration of attention modules within encoderdecoderarchitecturesfollowsseveralestablishedpatterns, eachofferingdistinctadvantagesforcontextmodelling.The choice of insertion points reflects fundamental trade-offs betweencomputationalefficiency,representationalcapacity, and the nature of contextual information being captured. Understandingthesedesignpatternsprovidesinsightinto howattentionmechanismscanbemosteffectivelydeployed forcontext-awareimageunderstanding.
3.1.1. Attention in the Bottleneck
Thebottleneckrepresentsthelowestspatialresolutionand highestsemanticabstractionpointintheencoder-decoder pipeline, typically occurring at the transition between encoding and decoding stages. Inserting attention mechanismsatthebottleneckenablesthenetworktomodel globalcontextandlong-rangedependenciesbeforespatial informationisprogressivelyrestoredduringdecoding.This design choice is particularly effective for tasks requiring understandingofoverallscenestructureandrelationships betweendistantimageregions.
The non-local neural network proposed by Wang et al. (2018) exemplifies attention at the bottleneck, inserting non-local blocks at intermediate network stages where featuremapshaverelativelylowspatialresolutionbuthigh channel dimensionality. At this resolution, the quadratic complexity of self-attention becomes computationally tractable,enablingthemodellingofrelationshipsbetweenall spatial positions. The non-local block computes pairwise affinitiesacrosstheentirespatialextent,generatingcontextaggregated features that incorporate information from all image regions. When placed at the bottleneck, this global contextual information can then be propagated to higher resolutionsthroughthedecoder,ensuringthatfine-grained predictionsbenefitfromholisticsceneunderstanding.
3.1.2. Attention in Skip Connections
Skip connections, which directly transmit high-resolution features from encoder layers to corresponding decoder layers,provideanalternativeandcomplementarylocation forattentioninsertion.TheU-Netarchitecture(Ronneberger et al., 2015) popularised this design, demonstrating that combininghigh-resolutionencoderfeatureswithupsampled decoder features through concatenation substantially improvessegmentationaccuracy,particularlyalongobject boundarieswherefinedetailisessential.
Thecombinationofbottleneckandskipconnectionattention representsaparticularlypowerfuldesignpattern,enabling networks to capture both global context through lowresolutionattentionandfinedetailthroughhigh-resolution attention.Thisdual-attentionapproachcharacterisesstateof-the-art architectures for dense prediction tasks, with OCRNet(Yuanetal.,2020)exemplifyingtheintegrationof object-region attention at multiple scales to achieve both contextualawarenessandpreciselocalisation.
3.2. Complexity Analysis
Thepracticaldeploymentofattentionmechanismsindeep learning systems requires careful consideration of computationalandmemorycomplexity.Differentattention designs exhibit vastly different scaling behaviours with respecttoinputresolutionandfeaturedimensionality,with important implications for both training feasibility and

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
inference efficiency. This section analyses the complexity characteristics of major attention families and discusses their practical implications for context-aware image understandingsystems.
3.2.1. Computational Complexity (FLOPs)
Floating-point operations (FLOPs) provide a hardwareindependent measure of computational cost that enables comparison across different attention mechanisms. The complexityanalysisrevealsfundamentaltrade-offsbetween representational power and computational efficiency that guidearchitecturaldesignchoices.
4. CHALLENGES IN IMPLEMENTATION
Despitethematurityofdeeplearningframeworksandthe richnessofthePythonecosystem,implementingattentionguidednetworkswithdynamicfeaturesuppressionpresents several significant challenges. These challenges span the entiredevelopmentlifecycle,frommodeldesignandtraining to deployment and optimisation, and require careful consideration to achieve both functional correctness and practicalefficiency.
4.1. Difficulty in Implementing Hard Attention
The implementation of hard attention mechanisms poses fundamental difficulties arising from the nondifferentiabilityofdiscreteselectionoperations.Unlikesoft attention,whichcomputescontinuousweightsthatcanbe propagated through standard backpropagation, hard attentionmakesdiscretedecisionsaboutwhichfeaturesto process forinstance,selectingasubsetofspatiallocations orbinarygatingofchannels.Thesediscretedecisionscreate adiscontinuityinthecomputationalgraph,preventingthe directapplicationofgradient-basedoptimisation.
As articulated in research on quantum hard attention mechanisms, the non-differentiability challenge has constrainedthewidespreadapplicabilityofhardattentionin deep learning. Various strategies have been developed to circumventthislimitation,eachwithitsowntrade-offs.The straight-through estimator, popularised by Bengio et al. (2013),approximatesthegradientofdiscretethresholding operations by treating them as identity functions during backward propagation. This approach enables end-to-end trainingbutintroducesbiasingradientestimationthatcan affectconvergence.TheGumbelsoftmaxrelaxationprovides a differentiable approximation to discrete sampling by replacingtheargmaxoperationwithasoftmaxoverGumbelperturbedlogits,withatemperatureparametercontrolling the sharpness of the approximation. During training, the temperaturecanbeannealedtoapproachdiscretedecisions at inference time while maintaining differentiability throughoutoptimisation.
4.2. Training Stability of Transformers on Small Datasets
VisionTransformersandtheirvariantspresentsignificant trainingchallengeswhenappliedtosmallormedium-sized datasets, stemming from their reduced inductive bias comparedtoconvolutionalneuralnetworks.Asdocumented inrecentCVPRworkshopproceedings,ViTstypicallyrequire substantiallylargertrainingdatasetstolearnlocalfeature representationseffectively,withperformanceonImageNet1K-scale datasets often falling short of comparably-sized CNNswithoutextensivedataaugmentationorpre-training. Thefundamentalissueliesintheself-attentionmechanism's flexibility: while this flexibility enables the modelling of complexrelationships,italsomeansthenetworkmustlearn spatiallocalityandtranslationequivariancefromdatarather thanhavingthesepropertiesbuiltintothearchitecture.On small datasets, the statistical evidence for these inductive biases is insufficient, leading to overfitting and poor generalisation. The Multi-Gradient Image Transformer (MGiT) approach proposed by researchers addresses this challengethroughparalleltrainingwithacompactauxiliary ViTthatadaptivelyoptimisesthetargetnetwork'sweights, demonstrating that specialised training strategies can partiallycompensateforlimiteddata.
5. EVALUATION, DATASETS, AND BENCHMARKS
5.1. Standard Datasets
Theevaluationofattention-guideddeeplearningmodelsfor context-awareimageunderstandingreliesonacollectionof standardised datasets that have become established benchmarkswithinthecomputervisioncommunity.These datasetsarecarefullycuratedtoencompassdiversevisual scenes, rich annotations, and tasks that require genuine contextualreasoning,therebyprovidingmeaningfulgrounds forcomparingdifferentarchitecturalapproaches.
TheCommonObjectsinContext(COCO)datasetrepresents one of the most widely adopted benchmarks for contextawareimageunderstanding,encompassingobjectdetection, segmentation,andcaptioningtasks.AsGonzález-Chávezet al.(2023)note,COCOhasbecomeacornerstonedatasetfor imagecaptioningresearch,providingcomplexsceneswith multiple interacting objects that necessitate contextual reasoningforaccuratedescription.Thedatasetcontainsover 200,000imageswithdetailedannotationsincludingobject instancesegmentation,stuffsegmentation,andfivehumanwritten captions per image, enabling comprehensive evaluationofmodels'abilitytounderstandbothobjectsand their contextual relationships. The COCO-Stuff variant extends this with additional stuff category annotations, providing even richer contextual information for dense predictiontasks(Youetal.,2025)

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
5.2. Evaluation Metrics
The assessment of attention-guided models for contextaware understanding employs a multifaceted suite of evaluation metrics, each designed to capture different dimensionsofmodelperformance.Thesemetricsspantaskspecific accuracy measures, computational efficiency indicators,and,increasingly,metricsthatattempttoquantify thequalityofcontextualreasoningitself.
5.2.1. Task-Specific Metrics
Forobjectdetectiontasks,meanAveragePrecision(mAP) servesastheprimaryevaluationmetric,measuringboththe accuracy of object localisation and the confidence of classification.Asdefinedincomputationalframeworkssuch as MATLAB's Deep Learning Toolbox, mAP computes the average precision across different recall levels, with detections considered correct when the intersection over union(IoU)betweenpredictedandgroundtruthbounding boxesexceedsaspecifiedthreshold(MathWorks,2024).The threshold can be varied to assess localisation quality at differentlevelsofstrictness,withmAP@0.5andmAP@0.75 providing insights into coarse and fine localisation performancerespectively.Forcontext-awaremodels,mAP across object categories provides indirect evidence of contextual understanding: accurate detection of small or occluded objects often depends on contextual cues from surroundingsceneelements.
6. CONCLUSION
This review has systematically examined the landscape of attention-guided deep learning for context-aware image understanding, with particular focus on hierarchical architecturesanddynamicfeaturesuppressionmechanisms. TheevolutionfromsimplechannelattentioninSENet(Huet al.,2018)tosophisticatedhierarchicalTransformerssuchas SwinTransformer(Liuetal.,2021)demonstratestherapid progress in enabling networks to selectively emphasise relevant contextual information. Our analysis reveals that whileattentionmechanismshavematuredconsiderably encompassingsoftandhardattention,channelandspatial operations, andself-attentionvariants the integration of dynamic feature suppression remains comparatively underdeveloped.Themathematicalunificationofattentionsuppression blocks presented herein highlights the complementary nature of these mechanisms: attention selectssalientfeatureswhilesuppressionactivelyeliminates irrelevant information, yet most existing architectures implement them separately rather than within unified frameworks.
The critical evaluation of benchmarks raises important questions about whether current performance improvements reflect genuine advances in contextual understanding or merely better exploitation of dataset statistics.Liuetal.'s(2025)ContextAmbiguitybenchmark
highlightstheneedforevaluationprotocolsthattestmodels' abilitytorecognisecontextualinsufficiency,movingbeyond forced-choice accuracy metrics. For dynamic suppression models,interpretabilitytechniquessuchasthoseproposed by Ren et al. (2024) offer pathways to verify that suppression targets genuinely irrelevant features rather than discarding useful information. The implementation challengesdocumentedinPythonframeworksunderscore thegapbetweenresearchprototypesandproduction-ready systems, particularly regarding hard attention training stabilityanddynamicgraphoptimisation.Futureresearch mustthereforepursueunifiedhierarchicalframeworksthat jointly optimise attention and suppression, designed with bothrepresentationalpoweranddeploymentefficiencyas primaryobjectives.
6.1. Limitations
This review, while comprehensive, acknowledges several limitationsthatshouldbeconsideredwheninterpretingits findings.First,therapidevolutionofattentionmechanisms meansthatrecentdevelopmentsemergingduringthereview processmaynotbefullyrepresented,particularlyregarding foundation models and large-scale vision-language pretrainingthatincreasinglysubsumeexplicitattentiondesign within broader architectures. Second, the focus on hierarchicalattentionanddynamicsuppressionnecessarily excludes related paradigms such as neural architecture search and automated attention design, which may offer complementaryinsightsforcontext-awareunderstanding. Third,theanalysisofbenchmarklimitations,whilecritical, does not propose new evaluation protocols but rather synthesises existing critiques, leaving the development of improvedbenchmarkstofuturework.Fourth,thediscussion of Python implementation challenges reflects current framework capabilities, which continue to evolve rapidly, potentially rendering specific optimisation difficulties transient. Finally, the review's emphasis on architectural patterns may underrepresent the importance of training methodologies,datacurationstrategies,andregularisation techniques that often prove as crucial as architectural choicesforachievingrobustcontext-awareunderstanding.
REFERENCES
1. Anderson,P.,He,X.,Buehler,C.,Teney,D.,Johnson,M., Gould, S. and Zhang, L. (2018) 'Bottom-up and topdown attention for image captioning and visual question answering',IEEE Conference on Computer VisionandPatternRecognition,pp.6077-6086.
2. Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick,C.L.andParikh,D.(2015)'VQA:Visualquestion answering',International Conference on Computer Vision,pp.2425-2433.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
3. Ba, J., Mnih, V. and Kavukcuoglu, K. (2015) 'Multiple objectrecognitionwithvisualattention',International ConferenceonLearningRepresentations.
4. Badrinarayanan,V., Kendall,A.andCipolla,R.(2017) 'SegNet: A deep convolutional encoder-decoder architecture for image segmentation',IEEE Transactions on Pattern Analysis and Machine Intelligence,39(12),pp.2481-2495.
5. Bahdanau, D., Cho, K. and Bengio, Y. (2015) 'Neural machine translation by jointly learning to align and translate',International Conference on Learning Representations.
6. Bengio, Y., Léonard, N. and Courville, A. (2013) 'Estimatingorpropagatinggradientsthroughstochastic neurons for conditional computation',arXiv preprint arXiv:1308.3432.
7. Byeon,W.,Breuel,T.M.,Raue,F.andLiwicki,M.(2015) 'Scene labeling with LSTM recurrent neural networks',IEEE Conference on Computer Vision and PatternRecognition,pp.3547-3555.
8. Chen,L.C.,Papandreou,G.,Kokkinos,I.,Murphy,K.and Yuille,A.L.(2015)'Semanticimagesegmentationwith deep convolutional nets and fully connected CRFs',International Conference on Learning Representations.
9. Chen,L.C.,Papandreou,G.,Kokkinos,I.,Murphy,K.and Yuille, A.L. (2018) 'DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs',IEEE Transactions on Pattern Analysis and Machine Intelligence,40(4),pp.834-848.
10. Chen, L.C., Papandreou, G., Schroff, F. and Adam, H. (2017) 'Rethinking atrous convolution for semantic imagesegmentation',arXivpreprintarXiv:1706.05587.
11. Chen,M.,Wu,J.,Wang,L.andZhang,X.(2021)'Learning to suppress significant activation values for robust image classification',IEEE Transactions on Image Processing,30,pp.7892-7905.
12. Child,R.,Gray,S.,Radford,A.andSutskever,I.(2019) 'Generating long sequences with sparse transformers',arXivpreprintarXiv:1904.10509.
13. Di Bello, F., Ben Hadj Hassen, S., Astrand, E. and Ben Hamed, S. (2021) 'A unified neurocognitive model of proactive and reactive attentional suppression',Neuroscience&BiobehavioralReviews, 129,pp.345-358.
14. Dosovitskiy,A.,Beyer,L.,Kolesnikov,A.,Weissenborn, D.,Zhai,X.,Unterthiner,T.,Dehghani,M.,Minderer,M.,
Heigold,G.,Gelly,S.,Uszkoreit,J.andHoulsby,N.(2021) 'An image is worth 16x16 words: Transformers for imagerecognitionatscale',InternationalConferenceon LearningRepresentations.
15. Dutta,A.,Mehrab,K.S.,Sawhney,M.,Neog,A.,Khurana, M., Fatemi, S., Pradhan, A., Maruf, M., Lourentzou, I., Daw, A. and Karpatne, A. (2025) 'Open world scene graphgenerationusingvisionlanguagemodels',arXiv preprintarXiv:2506.08189.
16. Farabet,C.,Couprie,C.,Najman,L.andLeCun,Y.(2013) 'Learninghierarchicalfeaturesforscenelabeling',IEEE Transactions on Pattern Analysis and Machine Intelligence,35(8),pp.1915-1929.
17. Figurnov,M.,Collins,M.D.,Zhu,Y.,Zhang,L.,Huang,J., Vetrov, D. and Salakhutdinov, R. (2017) 'Spatially adaptivecomputationtimeforresidualnetworks',IEEE Conference on Computer Vision and Pattern Recognition,pp.1039-1048.
18. Gao, X., Zhao, Y., Dudziak, Ł., Mullins, R. and Xu, C.Z. (2019) 'Dynamic channel pruning: Feature boosting andsuppression',InternationalConferenceonLearning Representations.
19. González-Chávez, O., Ruiz, G., Moctezuma, D. and Ramirez-delReal, T. (2023) 'Are metrics measuring whattheyshould?AnevaluationofImageCaptioning taskmetrics',SignalProcessing:ImageCommunication, 117,107071.
20. Google(2024) 'tf.compat.v1.metrics.mean_iou',TensorFlow Documentationv2.15.0.Available at:https://www.tensorflow.org/versions/r2.15/api_do cs/python/tf/compat/v1/metrics/mean_iou(Accessed: 19February2026).
21. Graham,B.,Engelcke,M.andvanderMaaten,L.(2018) '3D semantic segmentation with submanifold sparse convolutionalnetworks',IEEEConferenceonComputer VisionandPatternRecognition,pp.9224-9232.
22. Hassanin, M., Anwar, S., Radwan, I. and Khan, F.S. (2024) 'Visual attention methods in deep learning: A comprehensive survey',ACM Computing Surveys, 56(7),pp.1-42.
23. He, K., Zhang, X., Ren, S. and Sun, J. (2015) 'Spatial pyramid pooling in deep convolutional networks for visual recognition',IEEE Transactions on Pattern Analysis and Machine Intelligence, 37(9), pp. 19041916.
24. He, K., Zhang, X., Ren, S. and Sun, J. (2016) 'Deep residual learning for image recognition',IEEE

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Conference on Computer Vision and Pattern Recognition,pp.770-778.
25. Hu, J., Shen, L. and Sun, G. (2018) 'Squeeze-andexcitation networks',IEEE Conference on Computer VisionandPatternRecognition,pp.7132-7141.
26. Itti, L., Koch, C. and Niebur, E. (1998) 'A model of saliency-based visual attention for rapid scene analysis',IEEE Transactions on Pattern Analysis and MachineIntelligence,20(11),pp.1254-1259.
27. Jaderberg, M., Simonyan, K., Zisserman, A. and Kavukcuoglu, K. (2015) 'Spatial transformer networks',AdvancesinNeuralInformationProcessing Systems,28,pp.2017-2025.
28. Jang, E., Gu, S. and Poole, B. (2017) 'Categorical reparameterization with Gumbelsoftmax',International Conference on Learning Representations.
29. Jia,X.,DeBrabandere,B.,Tuytelaars,T.andGool,L.V. (2016)'Dynamicfilternetworks',AdvancesinNeural InformationProcessingSystems,29,pp.667-675.
30. Kim, J.H., Jun, J. and Zhang, B.T. (2018) 'Bilinear attention networks',Advances in Neural Information ProcessingSystems,31,pp.1564-1574.
31. Kitaev,N.,Kaiser,Ł.andLevskaya,A.(2020)'Reformer: Theefficienttransformer',InternationalConferenceon LearningRepresentations.
32. Krizhevsky, A., Sutskever, I. and Hinton, G.E. (2012) 'ImageNetclassificationwithdeepconvolutionalneural networks',AdvancesinNeuralInformationProcessing Systems,25,pp.1097-1105.
33. Lawson, D., Qureshi, A.S., Chung, W.H. and Bengio, Y. (2018) 'Learning hard attention for visual question answering with discrete optimization',NeurIPS WorkshoponRelationalRepresentationLearning.
34. Lazebnik, S., Schmid, C. and Ponce, J. (2006) 'Beyond bags of features: Spatial pyramid matching for recognizingnaturalscenecategories',IEEEConference onComputerVisionandPatternRecognition,pp.21692178.
35. Lefaudeux, B., Massa, F., Liskovich, D., Xiong, W., Caggiano,V.,Naren,S.,Xu,M.,Hu,J.,Tintore,M.,Zhang, S.andLeGendre,C.(2022)'xformers:Amodularand hackable Transformer modelling library',GitHub repository.Available at:https://github.com/facebookresearch/xformers
36. Li, H., Xiong, P., An, J. and Wang, L. (2018) 'Pyramid attention network for semantic segmentation',arXiv preprintarXiv:1805.10180.
37. Li, X., Wang, W., Hu, X. and Yang, J. (2019) 'Selective kernelnetworks',IEEEConferenceonComputerVision andPatternRecognition,pp.510-519.
38. Lin,J.,Rao,Y.,Lu,J.andZhou,J.(2017)'Runtimeneural pruning',Advances in Neural Information Processing Systems,30,pp.2181-2191.
39. Liu,J.,Wang,Y.,Zhang,L.andChen,T.(2025)'Detecting multimodal situations with insufficient context and abstaining from baseless predictions',arXiv preprint arXiv:2405.11145.
40. Liu,Z.,Lin,Y.,Cao,Y.,Hu,H.,Wei,Y.,Zhang,Z.,Lin,S. and Guo, B. (2021) 'Swin Transformer: Hierarchical vision transformer using shifted windows',International Conference on Computer Vision,pp.10012-10022.
41. Long, J., Shelhamer, E. and Darrell, T. (2015) 'Fully convolutional networks for semantic segmentation',IEEE Conference on Computer Vision andPatternRecognition,pp.3431-3440.
42. Luo, W., Li, Y., Urtasun, R. and Zemel, R. (2016) 'Understanding the effective receptive field in deep convolutional neural networks',Advances in Neural InformationProcessingSystems,29,pp.4898-4906.
43. MathWorks (2024) 'mAPObjectDetectionMetric',MATLAB Deep Learning Toolbox Documentation. Available at:https://uk.mathworks.com/help/vision/ref/mapobj ectdetectionmetric.html(Accessed:19February2026).
44. Mnih, V., Heess, N., Graves, A. and Kavukcuoglu, K. (2014)'Recurrentmodelsofvisualattention',Advances in Neural Information Processing Systems, 27, pp. 2204-2212.
45. NatureScientificReports(2024)'Table3:Quantitative comparison of unsupervised methods on Cityscapes, ADE20K,andCOCO-Stuff',NatureScientificReports.
46. Nie,M.,Sun,J.,Guoyang,H.,Niu,A.,Hu,Y.,Yan,Q.and Zhu,Y.(2025)'FSCFNet:Lightweightneuralnetworks viamulti-dimensional importance-aware optimization',Neurocomputing,131823.
47. Oktay,O.,Schlemper,J.,Folgoc,L.L.,Lee,M.,Heinrich, M.,Misawa,K.,Mori,K.,McDonagh,S.,Hammerla,N.Y., Kainz,B.,Glocker,B.andRueckert,D.(2018)'Attention U-Net: Learning where to look for the pancreas',MedicalImagingwithDeepLearning.
2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008 Certified Journal | Page621

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
48. Oliva,A.andTorralba,A.(2007)'Theroleofcontextin object recognition',Trends in Cognitive Sciences, 11(12),pp.520-527.
49. OpenBayes(2024)'Understandinginhibitionthrough maximallytenseimages',OpenBayesTrends.Available at:https://trends.openbayes.com/paper/understandin g-inhibition-through-maximally(Accessed:19February 2026).
50. Ren,S.,He,K.,Girshick,R.andSun,J.(2017)'FasterRCNN:Towardsreal-timeobjectdetection with region proposal networks',IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6), pp. 11371149.
51. Ren, Z.-X., Li, Y., Zhang, H. and Wang, J. (2024) 'Understanding inhibition through maximally tense images',arXivpreprintarXiv:2406.08924.
52. Ronneberger,O.,Fischer,P.andBrox,T.(2015)'U-Net: Convolutional networks for biomedical image segmentation',International Conference on Medical ImageComputingandComputer-AssistedIntervention, pp.234-241.
53. Russakovsky,O.,Deng,J.,Su,H.,Krause,J.,Satheesh,S., Ma,S.,Huang,Z.,Karpathy,A.,Khosla,A.,Bernstein,M., Berg,A.C.andFei-Fei,L.(2015)'ImageNetlargescale visual recognition challenge',International Journal of ComputerVision,115(3),pp.211-252.
54. Sigurdson,P.(2024)'DifferencesbetweenPyTorchand TensorFlowinAImodeldevelopment',Coda.Available at:https://coda.io/@peter-sigurdson/differencesbetween-pytorch-and-tensorflow-in-ai-modeldevelopme.
55. Simonyan, K. and Zisserman, A. (2015) 'Very deep convolutional networks for large-scale image recognition',International Conference on Learning Representations.
56. SingaporeUniversityofTechnologyandDesign(2025) 'Training indoor and scene-specific semantic segmentationmodels',IEEEXplore.
57. Sun,K.,Xiao,B.,Liu,D.andWang,J.(2019)'Deephighresolution representation learning for human pose estimation',IEEEConferenceonComputerVisionand PatternRecognition,pp.5693-5703.
58. Touvron,H.,Cord,M.,Douze,M.,Massa,F.,Sablayrolles, A.andJégou,H.(2021)'Trainingdata-efficientimage transformers&distillationthrough attention',International Conference on Machine Learning,pp.10347-10357.
59. University of Queensland (2024) 'Comparing ML libraries',HYPPODocumentation.Available at:https://hpouq.gitlab.io/hyppo/architecture/compare.html
60. Vaswani,A.,Shazeer,N.,Parmar,N.,Uszkoreit,J.,Jones, L., Gomez, A.N., Kaiser, Ł. and Polosukhin, I. (2017) 'Attention is all you need',Advances in Neural InformationProcessingSystems,30,pp.5998-6008.
61. Vinyals,O.,Toshev,A.,Bengio,S.andErhan,D.(2015) 'Showandtell:Aneuralimagecaptiongenerator',IEEE Conference on Computer Vision and Pattern Recognition,pp.3156-3164.
62. Visin, F., Ciccone, M., Romero, A., Kastner, K., Cho, K., Bengio,Y.,Matteucci,M.andCourville,A.(2016)'Reseg: Arecurrentneuralnetwork-basedmodelforsemantic segmentation',CVPRWorkshoponDeepLearningfor SemanticSegmentation.
63. Wang,J.,Sun,K.,Cheng,T.,Jiang,B.,Deng,C.,Zhao,Y., Liu, D., Mu, Y., Tan, M., Wang, X., Liu, W. and Xiao, B. (2020)'Deephigh-resolutionrepresentationlearning for visual recognition',IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(10), pp. 33493364.
64. Wang,X.,Girshick,R.,Gupta,A.andHe,K.(2018)'Nonlocalneuralnetworks',IEEEConferenceonComputer VisionandPatternRecognition,pp.7794-7803.
65. Wightman, R. (2025) 'timm: PyTorch image models',GitHub repository. Available at:https://github.com/huggingface/pytorch-imagemodels.
66. Woo,S.,Park,J.,Lee,J.Y.andKweon,I.S.(2018)'CBAM: Convolutional block attention module',European ConferenceonComputerVision,pp.3-19.
67. Wu, D., Li, Z. and Mitra, T. (2025) 'Inkstream: Instantaneous GNN inference on dynamic graphs via incremental update',IEEE International Parallel and DistributedProcessingSymposium.
68. Wu,H.,Xiao,B.,Codella,N.,Liu,M.,Dai,X.,Yuan,L.and Zhang, L. (2021) 'CvT: Introducing convolutions to vision transformers',International Conference on ComputerVision,pp.22-31.
69. Xie,E.,Wang,W.,Yu,Z.,Anandkumar,A.,Alvarez,J.M. and Luo, P. (2021) 'SegFormer: Simple and efficient design for semantic segmentation with transformers',Advances in Neural Information ProcessingSystems,34,pp.12077-12090.
70. Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R. and Bengio, Y. (2015)

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
'Show,attendandtell:Neuralimagecaptiongeneration with visual attention',International Conference on MachineLearning,pp.2048-2057.
71. Yang, B., Bender, G., Le, Q.V. and Ngiam, J. (2019) 'CondConv:Conditionallyparameterizedconvolutions forefficientinference',AdvancesinNeuralInformation ProcessingSystems,32,pp.1307-1318.
72. Yang, Z., He, X., Gao, J., Deng, L. and Smola, A. (2016) 'Stacked attention networks for image question answering',IEEEConferenceonComputerVisionand PatternRecognition,pp.21-29.
73. You, Z., Wang, J., Kong, L., He, B. and Wu, Z. (2025) 'Pix2Cap-COCO: Advancing visual comprehension via pixel-level captioning',arXiv preprint arXiv:2501.13893.
74. Yuan,Y.,Chen,X.andWang,J.(2020)'Object-contextual representationsforsemanticsegmentation',European ConferenceonComputerVision,pp.173-190.
75. Zhang,H.,Niu,Y.andChang,S.F.(2019)'Hierarchical attention networks for image captioning',AAAI ConferenceonArtificialIntelligence,33,pp.9253-9260.
76. Zhang,H.,Li,F.,Liu,S.,Zhang,L.,Su,H.,Zhu,J.,Ni,L.M. andShum,H.Y.(2025)'Optimisingvisiontransformer performance on limited datasets: A multi-gradient approach',IEEE/CVFConferenceon ComputerVision andPatternRecognitionWorkshops.
77. Zhao, H., Shi, J., Qi, X., Wang, X. and Jia, J. (2017) 'Pyramidsceneparsingnetwork',IEEEConferenceon Computer Vision and Pattern Recognition, pp. 28812890.
78. Zheng,S.,Lu,J.,Zhao,H.,Zhu,X.,Luo,Z.,Wang,Y.,Fu,Y., Feng, J., Xiang, T., Torr, P.H. and Zhang, L. (2021) 'Rethinkingsemanticsegmentationfromasequence-tosequence perspective with transformers',IEEE Conference on Computer Vision and Pattern Recognition,pp.6881-6890.