
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Mohammed Faisal Shahzad Siddiqui Department ofArtificialIntelligence,AnuragUniversity,Hyderabad,Telangana,India
Abstract - Deep neural networks are known to learn highly structured internal representations, yet empirical studies suggest that the effective dimensionality of hiddenlayer activations often collapses during training. Such representation rank collapse can limit feature diversity and adversely affect generalization. In this paper, we present a geometric analysis of representation rank in deep neural networks by modeling hidden activations as highdimensional matrices and examining their rank properties. We characterize rank collapse as a structural tendency induced by optimization dynamics and nonlinear activations. To mitigate this effect, we propose a simple rank-preserving regularization objective based on the logdeterminant of the activation Gram matrix, which explicitly encourages diverse feature representations without modifying network architectures. Our analysis highlights the importance of representation geometry in representationlearningandsuggestsrank-awareobjectives asapromisingdirectionforfutureresearch.
Key Words: DeepNeuralNetworks,Representation Learning, Rank Collapse, Log-Determinant Regularization, Feature Degeneracy, Deep Learning Theory
Deep neural networks owe much of their success to the abilitytolearnrichinternalrepresentationsfromdata[1], [2]. Increasing network depth and width is commonly associated with improved expressivity, under the assumption that additional neurons enable the model to capture a broader range of feature variations. This assumption implicitly relies on the idea that learned representations fully exploit the available dimensionality ofthenetwork.
However, accumulating empirical observations suggest thatthisisnotalways thecase.Recenttheoretical studies have analyzed the dynamics of representation learning in deep networks [3]. Despite large hidden layers, neural networks often learn representations that lie in relatively low-dimensional subspaces, with many neurons encoding highly correlated information. Such redundancy raises fundamental questions about how representational capacityisutilizedduringtraining.
From a linear algebraic viewpoint, hidden-layer activationscanbeviewedasmatriceswhoserank reflects theeffectivedimensionality oflearnedfeatures.Whenthe rank of these activation matrices is substantially lower
than the layer width, the network’s representational capacityisunderutilized. Werefertothisphenomenon as representationrankcollapse.
Inthispaper,weadoptageometricperspectivetoanalyze representationrankcollapseindeepneuralnetworks.We examine how training dynamics and activation functions contribute to the emergence of low-rank representations and propose a rank-preserving regularization term that encourages diversity among hidden units without modifyingnetworkarchitectures.
Representationlearningindeepneuralnetworkshasbeen studied from multiple perspectives, including expressivity, generalization,andinformationflow.Neuralnetworkswith sufficient width are known to be universal approximators [4]. More recent work has explored the dynamics of learningindeeparchitectures[3].
Recentstudieshavehighlightedthephenomenonofneural collapseduringtraining[5].Empiricalfindingssuggestthat hidden-layer activations often concentrate in lowdimensional manifolds, indicating reduced effective dimensionality.
Regularization techniques aimed at improving representation quality include orthogonality constraints, decorrelation-based objectives, and diversity-promoting penalties. However, these approaches are often heuristic anddonotexplicitlyanalyzerankstructure.
Incontrast,thisworkfocusesonrepresentationrankasan explicit and measurable structural property of learned representations, offering a geometric perspective without requiringarchitecturalmodifications.
Consider a neural network layer with d hidden units. Given a batch of n input samples, the corresponding hidden activations can be represented as a matrix
Each row corresponds to a sample, while each column representsaneuron’sresponse.
TherankofHreflectsthenumberoflinearlyindependent featuredirectionscapturedbythelayer.Afull-rankmatrix indicates diverse representations, whereas a low-rank matriximpliesredundancyamongneurons.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Theoveralltrainingobjectivebecomes:
Empirical observationssuggestthatthe rank of activation matricesoftendecreasesasnetworkdepthincreases.This phenomenon, referred to as representation rank collapse, indicates that effective dimensionality is significantly lowerthannominallayerwidth.

Fig-1: Illustrationofrepresentationrankcollapseindeep neuralnetworks.
This behavior arises due to optimization dynamics and nonlinear activation functions. Gradient-based optimization often favors low-dimensional solutions, whileactivationfunctionssuchasReLUpromotesparsity, leadingtocorrelatedneuronresponses.
To mitigate rank collapse, we propose a rank-preserving regularizationstrategy.
GiventheactivationmatrixH,wedefinetheGrammatrixG = H^T H, which captures pairwise similarities between neuronactivations.
Weintroducethefollowingregularizationterm: ( ) whereε>0ensuresnumericalstability.
Maximizing the determinant increases the volume spanned by activation vectors, encouraging diverse feature representations. Since training minimizes loss, we minimizethenegativelog-determinant.
where λ controls the strength of regularization. This approachoperatesonlyduringtrainingandintroducesno additionalinference-timecost.

Fig-2: Trainingpipelinewithrank-preserving regularizationappliedtohiddenlayers.
To evaluate the effect of rank-preserving regularization, we conduct experiments using standard supervised learning benchmarksandcommonlyused neural network architectures. Our goal is not to achieve state-of-the-art performance, but to analyze how rank-aware objectives influencelearnedrepresentations.
We consider baseline models trained using the task loss alone and compare them against identical architectures trained with the proposed rank regularization. Regularization is applied to selected hidden layers, while allothertrainingsettingsarekeptidenticaltoensureafair comparison.
Model performance is evaluated using classification accuracy. To assess representational properties, we measure the rank of hidden-layer activation matrices during training and visualize learned features using principalcomponentanalysis.
All experiments are conducted using the same optimization parameters, including learning rate, batch size,andnumberoftrainingepochs.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Experimental observationsindicate that networkstrained with rank-preserving regularization consistentlymaintain higher activation rank acrosslayerscompared to baseline models. In particular, deeper layers exhibit reduced representationcollapsewhentheproposedregularization isapplied.
Visualization of hidden representations further supports this observation. Baseline models tend to produce tightly clustered feature embeddings, whereas regularized models learn more dispersed representations that span a largersubspace.
Importantly, encouraging higher-rank representations does not negatively impact predictive performance. In several cases, regularized models achieve comparable or slightlyimprovedaccuracyrelativetobaselinemodels.
Overall, the results demonstrate that representation rank collapse can be mitigated through simple regularization objectives.
This study focuses on moderate-sized networks and standard benchmark datasets. Further investigation is required to evaluate its behavior in very deep architecturesandlarge-scaletrainingregimes.
Additionally, the relationship between representation rank and generalization performance is not fully understoodandremainsanopenproblem.
Wepresentedageometricanalysisofrepresentationrank collapse in deep neural networks. By analyzing hidden activations using linear algebra, we characterized rank collapseasastructurallimitation.
We proposed a rank-preserving regularization strategy that encourages higher-dimensional feature representations without modifying architectures. Experimental observations demonstrate improved representationdiversitywhilemaintainingperformance.
These findings highlight the importance of geometric considerations in representation learning and suggest promisingdirectionsforfutureresearch.
[1] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature,vol.521,pp.436–444,2015.
[2] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning,MITPress,2016.
[3] A. Saxe, J. McClelland, and S. Ganguli, “Exact solutions to the nonlinear dynamics of learning in deep linear neural networks,” International Conference on LearningRepresentations(ICLR),2014.
[4] K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” NeuralNetworks,vol.2,no.5,pp.359–366,1989.
[5] V. Papyan, X. Han, and D. Donoho, “Prevalence of neural collapse during the terminal phase of deep learning training,” Proceedings of the National Academy of Sciences (PNAS), vol. 117, no. 40, pp. 24652–24663,2020.
[6] C. Bishop, Pattern Recognition and Machine Learning, Springer,2006.
[7] S. Arora, N. Cohen, W. Hu, and Y. Luo, “Implicit regularizationindeepmatrixfactorization,”Advances in Neural Information Processing Systems (NeurIPS), 2019.