
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Mr.V.Murali Krishna1 , U.Chaitanya lakshmi2 , S. Jessy3 , S.Rohith Sai 4 , V.Hasanthi5
1 Assistant Professor,Dept. of Electronics and Communication Engineering Department, Seshadri Rao Gudlavalleru Engineering College, Andhra Pradesh, India 2345 Student, Department of Electronics and Communication Engineering, Seshadri Rao Gudlavalleru Engineering college, Andhra Pradesh, India
Abstract - In modern digital systems, especially in image processing and digital signal processing (DSP) applications, achieving high performance with low power consumption is crucial. Approximate computing has emerged as a promising paradigm to optimize power, area, and delay by allowing controlled inaccuracies in computations where exact results are not critical. This paper proposes the design and implementation of a 64-bit approximate multiplier using high- order compressors such as 8:2, 4:2, and 3:2 to reduce the number of partial product reduction stages. The use of higher- order compressors significantly lowers the critical path delay and power consumption by minimizing the number of intermediate additions required in the reduction tree. The proposed 64- bit multiplier is evaluated and compared against conventional exact multipliers in terms of delay, power, and accuracy metrics. Simulation results demonstrate a substantial improvement in delay and power efficiency, making the design well-suited for image processing tasks and DSP operations where approximate results are acceptable.
Keywords: Approximate Computing, Approximate Multiplier, High-Order Compressors, Partial Product Reduction, Low-Power VLSI Design.
Approximatecomputingisanemergingconceptindigital designthatrelaxestherequirementofexactcomputationtoachieve improvementsin powerconsumption, speed, andhardware efficiency. This approachis especiallyuseful forembedded and mobile systems that operate under strict energy and performance constraints. Many modern applications such as multimedia processing, image processing, machine learning, and data mining can tolerate small errors, meaning that a perfectlyaccurateresultisnotalwaysnecessary.Multipliersarefundamentalcomponents inmicroprocessors,digitalsignal processors, and embedded systems where they are used in operations such as filtering, convolution, and neural network computations.However,multipliersarecomplexcircuitsandconsumesignificantpowerandhardwareresourcebecause ofthis,thedesignofApproximatemultipliershasbecomeanimportantresearch areatoimprovesystemperformanceand reduceenergyconsumption.Intheproposedmultiplierdesign,exactcompressorsareusedinspecificpartsofthecircuit to maintain computational accuracy. While approximatecomputing reduces power consumption and delay by allowingsmallerrors,usingexactcompressorsincritical sectionsofthe multiplierensuresthattheoverallerrordoes not significantlyaffectthefinaloutput.Inparticular,exact compressorsaretypicallyusedinthemostsignificant(MSB)region of the partial product reduction stage, where errors would have a larger impact on the final result. By combining approximate techniques with exact compressors, the design achieves a balance between performance and accuracy. The approximate compressors help reduce delay, power consumption, and hardware complexity, while the exact compressors preserve the reliability of important computations. This hybrid approach improves the efficiency of the multiplier while still maintaining acceptable output quality for applications such as digital signal processing and machine learning. In our project, the partial product reduction stage is implemented using 8:2, 4:2, and 3:2 compressors. These compressors help combine multiple input bits into fewer output bits, thereby reducing the number of reduction stages and improving the overall speed of the multiplier. The use of these compressors creates a balanced reduction tree that decreases propagation delay and improves power efficiency while maintaining acceptable accuracy, making the design suitableforapplicationssuchasdigitalsignalprocessingandartificialintelligencesystems.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Narayanamoorthy, S., Moghaddam, H.A.,Liu,Z.,Park,T.,& Kim, N. S. (2015). Energy- Efficient Approximate Multiplication for Digital Signal Processing and Classification Applications. IEEE Transactions on Very Large Scale Integration (VLSI) Systems,23(6),1180–1184.Theneedtosupportvariousdigitalsignalprocessing(DSP)and classificationapplicationson energy-constrained devices has steadily grown. Such applications often extensively perform matrix multiplications using fixed-pointarithmeticwhileexhibitingtoleranceforsomecomputationalerrors.Hence,improvingtheenergyefficiencyof multiplications is critical. In this brief, we propose multiplier architectures that can tradeoff computational accuracy with energy consumption at design time. Compared with a precise multiplier, the proposed multiplier can consume less energy/opwithaveragecomputationalerrorof◻1%.Finally,itisdemonstratedthatsuchasmallcomputationalerrordoes nonotably impact the quality of DSP and the accuracy of classification applications. Summary: In this paper, Power consumptionandareacanfurtherbereduced.Zervakis
G.,Xydis,S.,Tsoumanis,K.,Soudris,D.,&Pekmestzi,K.(2015).Hybridapproximatemultiplierarchitecturesforimproved power-accuracy trade-offs. 2015 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED). Approximatecomputingformsapromising designalternativeforinherentlyerrorresilientapplications, tradingaccuracyfor power savings. In this paper, we exploit multi-level approximation,i.e. at the algorithmic, the logic and the circuit level, to design low power approximate arithmetic architectures for hardware multipliers. Motivated from the limited power savings that approximation techniques can achieve in isolation, we explore hybrid methods that apply simultaneously more than one techniques from different layers. We introduce the concept of perforation for approximate arithmetic circuit design and we explore the newly defined design space of hybrid designs showing that it leads to lower power consumption at every examined error range. To address the increased complexity of the target design space, we introduce an heuristic optimization technique and the corresponding design framework that automatically generates hybridlow-powerapproximatemultipliersrequiringasmallnumberofdesignevaluations,i.e.synthesis,simulationpower andtiminganalysis Throughextensiveexperimentation,weshowthattheproposedtechniquesconvergetowardsoptimal solutions and deliver approximate designs that are always more efficient with respect to state-of-art approaches. Power savings of 11% are reported for small error bounds and more than 30% in case of more relaxed error constraints. Summary:Inthispaper,powerconsumptioncanfurtherbereduced.
A. Momeni, J. Han,“Design and Analysis of Approximate Compressors for Multiplication” IEEETransactions on Computers. Inexact (or approximate) computing is an attractive paradigm for digital processing at nanometric scales. Inexact computing is particularly interesting for computer arithmetic designs. This paper deals with the analysis and design of two new approximate 4-2 compressors for utilization in a multiplier. These designs rely on different features of compression, such that imprecision in computation (as measured by the error rate and the so-called normalized error distance) can meet with respect to circuit- based figures of merit of a design (number of transistors, delay and power consumption). Fourdifferentschemesfor utilizing the proposed approximate compressors are proposed and analyzed for a Dadda multiplier. Extensive simulation results are provided and an application of the approximate multipliers to image processing is presented. The results show that the proposed designs accomplish significant reductions in power dissipation, delay and transistor count compared to an exact design; moreover, two of the proposed multiplier designs provide excellent capabilitiesforimagemultiplicationwithrespecttoaverage normalized error distance and peak signalto-noiseratioSummary:Inthispaper,highspeedisachieved,buttransistorcountismore
The proposed method focuses on the design and implementation of a 64-bit approximate multiplier and 64 bit Exact multiplierthatutilizesacombinationofhigh-ordercompressors specifically8:2,4:2,and3:2compressors toimprove speedandpowerefficiencyatthecostofasmallreductioninaccuracy.Thisdesignisparticularlysuitedforerror-tolerant applications such as image processing and digital signal processing (DSP), where a trade-off between accuracy and hardwareefficiencyisacceptable.
1. Motivation Traditional multiplier architectures for higher-bit computations involve numerous addition stages to compress the partial products, resulting in increased critical path delay, power consumption, and silicon area. The use of approximate computing, combined with high-order compressors, offers a solution by simplifying the hardware and

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
reducingthenumberofintermediatestages.
2. Compressor Hierarchy and Selection The core idea is to reduce the number of partial product reduction stages by using high-order compressors capable of handling more input bits per compression unit. The proposed hierarchy includes:
• 8:2 Compressor: Takes8inputbitsandproduces2 outputs(sumandcarry),significantlyreducingthedepthof the reductiontree.Itisusedinearlyreductionstageswheremorepartialproductsareavailable.
• 4:2 Compressor: Usedinthemiddlestagestofurthercompressthealreadyreducedpartialproducts.
• 3:2 Compressor (Full Adder): Utilized in the final reduction stages to prepare inputs for the final adder. These compressorsare designedusing approximate logicprinciples,where somecarry outputsor sumbits are simplifiedto reducelogiccomplexity,resultinginlowerpowerandfasterpropagationdelay
3. Partial Product Generation The partial products are generatedusingtheANDgatearraymethod, whereeach bit of themultiplicandisANDedwitheachbitofthemultiplier,generatingatotalof64×64=4096bits.Thesearearrangedina matrixformforreduction.
4. Partial Product Reduction usingCompressorsInsteadof usingtraditionalCarrySaveAddersorjust3:2compressors,the partialproductsarefed through thefollowingstages:
Stage 1: 8:2compressorsareusedtoaggressivelyreducepartialproductsacrossrows.
Stage 2: Theoutputsof8:2compressorsarefurthercompressedusing4:2compressors.
Stage 3: Remaining bits arereducedusing3:2compressorstoalign theresultintotwo rows(sumandcarry).This optimizedmulti-stagereductionstrategyminimizesthenumberofclockcyclesandcriticalpathdelay.
5. Final Addition Stage After reduction, the final two rows are added using a fast approximate adder, such as a CarryLookaheadAdder(CLA)oranApproximateCarrySelectAdder(ACSA),toobtainthefinalproduct.
1 16-bit multiplier with exact 8:2 compressor TheproposedexactXOR-MUX8:2compressorisusedtobuildthe16x16 multiplierasshowninbelowfigure.


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Presented in this this in a lower-order compressor with tolerable error and used in the construction of higher-order compressors , so the erroneous in the higher- order compressor is not calculated accurately. This work overcomes the above problem by creating the architecture comparing all the inputs to all the outputs for the accurate calculation of error that can be created by approximating in any part of the circuit. adder and lower order compressors The compressors are used in the multiplier for the tree reduction stage usually made up of full-adder. The full-adder is named as 3:2 compressors or counter is usually used for the construction for any higher-order compressor. One full-adder will be constructed with two XOR’s and one MUX is proposed so as reduce the area and power without any change in the truth table.
Themultiplierwith8:2compressorhasonlythreereductionstagesislessthanthe4:2compressorhasthereductionstage offour,therebyitisefficienttousehigher-ordercompressorifthemultiplierwidthisincreased.InFig.1thecolorcodeis usedtoidentifydifferentcomponents:pink–8:2compressors,red–4:2compressors,thickandlightblue–full-adders
2.16x16 multiplier with approximate 8:2 XOR mux based compressor With the approximation finder method, many inputs are matched with outputs for 75% of input combination. In our proposed approximate 8:2 compressor cin4 is bypassed to carry, soas to reduced area, power, delay without affecting the image quality. The approximate computing is themajorconcerninreducingthepower,area, and delay. Novel approximation technique is presented in this paper. The proposed 8:2 compressors consist of 13 inputs and having 213 = 8912 input combinations. The circuit consists of 7

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
output cout0 – cout4, sum, carry. The previous work in approximation is done in a lower-order compressor with tolerable error and used in the construction of higher-order compressors, so the erroneous in the higher-order compressor is not calculated accurately. This work overcomes the above problem by creating the architecture comparingall theinputs to all theoutputs for the accuratecalculation of error thatcan be created by approximating in any part of the circuit. The flow chart shown which demonstrates the correlation of every input to every output. The designofseveral15:4,9:4, 8:2compressorsareproposedrespectively.Allthesedesignsaredevelopedusingfull-adder and lower order compressors. The approximation feasibility overall this design is very low. The Proposed 8:2 compressor is designed with the straight forward approach with the parallel stream of input to output through XOR –MUXarchitectureasdemonstrated.Thecompressorsareusedinthemultiplierforthetreereductionstageusuallymade up of full-adder. The full-adder is named as 3:2 compressors or counter is usually used for the construction for any higher-ordercompressor.Onefull-adderwillbeconstructedwithtwoXOR’sandoneMUXisproposedsoasreducethe areaandpowerwithoutanychangeinthetruthtable.
3. The figure illustrates the compressor-tree based partial product reduction strcture used in the proposed multiplier architecture. It demonstrates how partial products are grouped column-wise and reduced using Exact 4:2 compressors,alongwithfulladdersandhalfadders, acrosssuccessivestages.Althoughasmallerstructureisshown for clarity, the same reduction principle is applied in the proposed 16×16 multiplier design. The multi-stage compression processreducesthenumberofpartialproductrowstotwofinaloperands,whicharethensummedusingacarry-propagate adder to obtain the final product. This approach minimizes the critical path delay while ensuring exact multiplication results.

4.Results:
TheobtainedRTLschematicsconfirmtheproperstructural realizationofthecompressor-treebasedreductionnetwork and final carry-propagate addition stage. demonstrating reduced logic complexity and improved hardware efficiency. This reduction is achieved by selectively employing approximate compressor units in less significant bit positions while preserving exact computation in the more significant bit regions, thereby maintaining acceptable output accuracy. Furthermore, the compressor-based reductionapproachshortensthecriticalpathbyminimizing thenumberofsequential addition stages, which contributes to higher operating speed. Overall, the experimental results indicate that the proposed multiplierarchitectureprovides afavorabletrade-offbetweenaccuracy,area,andperformance,makingitwellsuitedfor error-tolerant applications such as image processing, signal processing, and multimedia computing.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072




International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Figure 7: RTL Schematic for 64 bit Approximate Multiplier using higher order compressors
Table 1. Delay Analysis of Exact vs Approximate Multiplier Architectures
Multiplier Type Total Delay (ns) Logic Delay (ns) Net Delay (ns)
Multiplier
Multiplier
ns 7.44ns 21.09ns
ns 7.31ns
The proposed 64-bit multiplier design, implemented using 32-bit submodules and optimized with compressors/adders, demonstrates efficient architecture for high-speed arithmetic operations. By replacing conventional carry look- ahead adders with compressorbased structures, the design achieves fast partial product accumulation, reduced propagation delay, and improved Performance scalability. The modular approachusing32-bit multipliers allows easier synthesis, verification, and hardware realization on FPGA or ASIC platforms. This architecture is highly suitable for applications in digitalsignalprocessing,imageprocessing,andcryptographicsystemswherebothspeedandaccuracyarecritical.
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072 © 2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008 Certified Journal | Page1000
[1] Gorantla, A. and Deepa, P., 2019. Design of Approximate Subtractors and Dividers for Error Tolerant Image ProcessingApplications.JournalofElectronicTesting,pp.17.
[2] Kim, Y., Zhang, Y. and Li, P., 2014. Energy efficient approximate arithmetic for error resilient neuromorphic computing.IEEETransactionsonVeryLargeScaleIntegration(VLSI)Systems,23(11),pp.2733-2737
[3] Zhou, Y., Lin, J., Wang, J. and Wang,Z.,2018,October.Approximate Comparator: Design and Analysis. In 2018 IEEE InternationalWorkshoponSignalProcessingSystems(SiPS)(pp.1-5).IEEE.
[4] Monajati, M., Fakhraie, S.M. and Kabir, E., 2015. Approximate arithmetic for low-power image median filtering. Circuits,Systems,andSignalProcessing,34(10),pp.3191-3219

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056 Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
[5] Chang, C.H., Gu, J. and Zhang, M., 2004. Ultra low-voltage low-power CMOS 4-2 and 5-2 compressors for fast arithmeticcircuits.IEEETransactionsonCircuitsandSystemsI:RegularPapers,51(10),pp.1985-1997.
[6] Taheri, M., Arasteh,A., Mohammadyan, S., Panahi, A. and Navi, K., 2020. A novel majority based imprecise 4: 2 compressor with respect to the current and future VLSI industry. Microprocessors and Microsystems, 73, p.102962.
[7] Moaiyeri, M.H., Sabetzadeh, F.And Angizi, S., 2018. An efficient majority-based compressor for approximate computinginthenanoera.MicrosystemTechnologies,24(3),pp.1589-1601.
[8] Gorantla, A., 2017. Design ofapproximate compressors for multiplication. ACM Journal onEmerging Technologies inComputingSystems(JETC),13(3),pp.117.
[9] Marimuthu, R., Rezinold, Y.E. and Mallick, P.S., 2016. Design and analysis of multiplier using approximate 15-4 compressor.IEEEAccess,5,pp.1027-1036.
[10] Guo, Y., Sun, H., Guo, L. and Kimura, S., 2018, October. Low-cost approximate multiplier design using probability- driven inexact compressors. In 2018 IEEE Asia Pacific Conference onCircuits andSystems(APCCAS) (pp.291-294).IEEE.
[11] Marimuthu,R.,Bansal,D.,Balamurugan,S.andMallick, P.S.,2013.Designof8-4and9-4CompressorsForhighSpeed Multiplication.AmericanJournalofAppliedSciences,10(8),p.893.
© 2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008 Certified Journal | Page1001