Skip to main content

Next Hire: An AI-Powered Intelligent Recruitment Platform

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Next Hire: An AI-Powered Intelligent Recruitment Platform

Student, Dept. of Computer Science & Engineering, The Apollo University, Chittoor, Andhra Pradesh, India -

Abstract – Modernrecruitmentpipelinesareburdenedbymanualresumescreening,inconsistentcandidateevaluation, and a complete absence of behavioural insight at the pre-screening stage. These shortcomings result in poor hiring outcomes, qualifiedcandidatesbeingoverlooked,anda degradedapplicant experience. Thispaper presentsNextHire,an end-to-end AI-powered recruitment platform that automates and enhances the entire hiring lifecycle. The system integrates a multi-stage resume analysis pipeline comprising PyMuPDF document parsing, Qwen 2.5 7B large language model (LLM) structured extraction via Ollama, and sentence-transformer semantic embeddings, with an NLP-driven behaviouralassessmentmoduleemployingVADERsentimentanalysisandscikit-learn-basedpersonalitytraitinference.A configurable weighted composite scoring model combines a resume fit score (default 50%) and a behavioural score (default50%) toproduceanobjectivecandidateranking.Theplatformisbuilton a FastAPIPythonbackendanda React 18 TypeScript frontend, with MongoDB providing persistent storage. Evaluation results demonstrate average API response time under 200 ms, resume analysis pipeline time of 8–15 seconds on a warm LLM instance, and qualitative semanticskill-matchingaccuracyofapproximately90%on50testresumes.All16unittestsand allevaluatedintegration test cases passed. NextHire addresses critical gaps in existing recruitment technology, including keyword dependency, siloedassessment,opaquescoring,andminimalcandidatefeedback,whilepreservingdataprivacythroughlocallyhosted LLMinference.

Key Words: AI Recruitment, Natural Language Processing, Resume Analysis, Behavioural Assessment, Sentence Transformers, FastAPI, Large Language Models, VADER, MongoDB, React

1. INTRODUCTION

The global talent acquisition landscape has undergone significant disruption over the past decade. Industry surveys indicate that over 250 applications are received per corporate job posting on average, yet only a handful of candidates advance beyond the initial screening stage [1]. Traditional recruitment systems rely on manual resume review averagingsixsecondsperresume andkeyword-basedapplicanttrackingsystems(ATS)thatfailtointerpretsemantic equivalences between skill descriptions. The outcome is a process that is simultaneously resource-intensive and statisticallyunreliable.

Advances in natural language processing (NLP), transformer-based language models, and open-source machine learning libraries have created a practical pathway toward intelligent, automated recruitment pipelines. Pre-trained sentence embedding models such as Sentence-BERT (SBERT) [2] enable semantic similarity measurement far beyond keyword matching. Efficient open-source LLMs in the 7–13 billion parameter range [3] now enable structured information extractionfromunstructureddocumentswithoutrelianceoncostlycloudAPIs.

This paper presents NextHire, a full-stack web platform that integrates these AI capabilities into a cohesive recruitment solution. The system automates resume parsing and scoring, generates AI-driven situational behavioural questions, evaluatescandidateresponsesusingNLPtechniques,andprovidesrecruiterswithrankedapplicantlists,detailedreports, andananalyticsdashboard.

1.1

Problem Statement

Contemporaryrecruitmentsystemssufferfromseveralcriticaldeficiencies:

• Manual Resume Screening: Recruiters spend an average of six seconds per resume, making a fair evaluation of all applicantsstatisticallyimpossible.

• Keyword-OnlyMatching:MostATStoolsrelyonexactkeywordmatching,failingtorecognisesemanticequivalences.A candidate describing 'REST microservices' may be rejected by a system seeking 'API development' despite identical competencies.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

• Absence of Behavioural Insight: Personality traits and communication skills, which are strong predictors of job performance[8],arerarelyassessedatthepre-screeningstage.

• InconsistentEvaluation:Differentrecruitersapplydifferentcriteria,leadingtoinconsistencyinhiringdecisions.

• LimitedCandidateFeedback:Candidatesreceivenofeedbackpost-application,damagingtheemployerbrand.

1.2 Objectives

The primary objectives of Next Hire are: (1) to implement an end-to-end AI-powered recruitment platform covering the completehiringlifecycle;(2)todevelopasemanticresumeanalysispipelineusingsentence-transformerembeddingsand alocalLLM;(3)toimplementanNLP-basedbehaviouralassessmentmodule;(4)tocreateacompositecandidateranking model;(5)todesignarecruiteranalyticsdashboard;and(6)togeneratedetailedPDFcandidateassessmentreports.

2. LITERATURE REVIEW

2.1 Automated Resume

Screening

Early ATS tools employed keyword extraction and Boolean matching, resulting in significant false-negative rates [4]. TFIDF weighting improved relevance but could not capture semantic relationships. Word embedding models Word2Vec [5]andGloVe[6] enabledvector-spacesemanticmatchingbutstruggledwithcontextdependence.BERT[7]byDevlin etal.(2018)establishedbidirectionalcontextualencodingasthedominantparadigm.ReimersandGurevych[2]extended this with Sentence-BERT, producing semantically meaningful sentence embeddings suitable for large-scale resume matching.

2.2 Behavioural Assessment

SchmidtandHunter[8]demonstratedthroughmeta-analysisthatstructuredinterviewsandpersonalityassessmentsare among the strongest predictors of job performance. VADER [10], introduced by Hutto and Gilbert (2014), provides efficient rule-based sentiment analysis well-suited for informal text. Situational Judgment Tests (SJTs), validated by McDanieletal.[11],serveasthetheoreticalbasisforNextHire'sAI-generatedbehaviouralquestionformat.

2.3 Large Language Models

GPT-3 [12] demonstrated that LLMs could perform complex information extraction with minimal task-specific training. Kuzman et al. [3] showed that 7B–13B parameter open-source models achieve competitive accuracy on domain-specific structured extraction tasks. The Qwen 2.5 series [13] exhibits strong instruction-following and JSON-structured output capabilities.NextHireleveragesQwen2.57BviaOllamaforprivacy-preservinglocalinference.

2.4 Research Gap

Table-1 highlights the primary research gap: no existing platform simultaneously provides semantic resume matching, NLP-based behavioural assessment, transparent scoring, locally hosted LLM inference, and detailed candidate feedback withinasingleaccessiblesystem.

Table-1: Comparison of Existing Platforms

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Pymetrics Indirect Yes No Limited

NextHire SBERT VADER+ML Ollama FullPDF

3. PROPOSED METHODOLOGY

3.1 System Architecture

NextHire follows a three-tier client-server architecture. The presentation tier consists of a React 18 TypeScript SPA communicating with the backendvia AxiosHTTP with JWTinterceptors.Theapplicationtier isa Python FastAPIservice orchestrating all business logic, AI pipeline execution, authentication, and report generation. The data tier is MongoDB withanin-memoryfallbackfordevelopmentenvironments.

3.2 Resume Analysis Pipeline

Thepipelineoperatesinfoursequentialstages:

• DocumentParsing:PyMuPDFextractstextfromPDFresumes;python-docxhandlesDOCXformats.

• StructuredExtraction:CleanedtextissubmittedtoQwen2.57BviatheOllamaHTTPAPI,returningastructuredJSON ofname,education,experience,skills,andprojects.

• Semantic Embedding: Extracted skills are encoded as 384-dimensional vectors using the all-MiniLM-L6-v2sentencetransformermodel.

• Fit Score Computation: Cosine similarity is computed between each job requirement vector and all resume skill vectors. The fit score is the mean of maximum similarity values, normalised to [0, 1]. Threshold τ = 0.70 classifies a skillasmatched.

Thescoringformula: fit_score = mean(max_k cosine_sim(V_J[i], V_R[k]))

3.3 Behavioural Assessment Module

The module operates in three phases. In Phase 1, Qwen 2.5 7B generates five situational questions in the STAR format tailored tothe job description.InPhase 2,candidate responsesare validated for minimumtokencountanda copy-paste

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072 © 2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008 Certified Journal | Page 1435

Fig. 1: Next Hire System Architecture

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

similarity check (cosine similarity < 0.85 with the source question). In Phase 3, each response is processed through: (a) VADER sentiment scoring; (b) NLTK linguistic feature extraction; and (c) scikit-learn personality trait classification, producingBigFiveindicators.Theaggregatebehaviouralscoreisnormalisedto[0,1].Responsesfailingqualitychecksare cappedat0.12.

3.4 Composite Scoring Model

The combined score is: combined_score = (Wᵣ × fit_score) + (Wᵇ × behavioural_score) where Wᵣ + Wᵇ = 1.0 (default 0.50/0.50).Bothscoresarepre-normalisedto[0,1].Thecombinedscoreismultipliedby100forpercentagedisplay.

4. SYSTEM DESIGN & IMPLEMENTATION

4.1 Technology Stack

Table-2summarisesthecoretechnologystack.

Table-2: Technology Stack

Component Technology

Backend FastAPI+ Python3.10

Database MongoDB7+

LLMRuntime Ollama+Qwen 2.57B

Embeddings sentencetransformers

Sentiment VADER+NLTK

MLModels scikit-learn

DocParsing

PyMuPDF+ python-docx

PDFReports ReportLab

Auth PyJWT+bcrypt

Purpose

AsyncREST API,OpenAPI docs

Flexible document schema

Local,privacypreserving inference

384-dim semantic vectors

Behavioural NLPscoring

Personality trait classification

PDFandDOCX extraction

Structured candidate reports

StatefulJWT, RBAC,sessions

© 2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008 Certified Journal | Page 1436

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Frontend React18+ TypeScript SPA,Tailwind, TanStackQuery

Deployment

4.2 Authentication & Security

Docker Compose

Containerized multi-service

Authentication is implemented as a stateful JWT token pair system. Login issues a 15-minute access token and a 7-day refreshtokenstoredasa bcrypthashina sessionscollection.Aserver-sideinactivitytimeoutof180minutesinvalidates idlesessions.Allrole-specificendpointsenforceRBAC,returningHTTP403onrolemismatch.FrontendAxiosinterceptors handlesilenttokenrefreshon401responses.

4.3

REST API Design

The API exposes 18 endpoints under the /api/* prefix, following RFC 7231 HTTP semantics. Key endpoints: POST /api/resume/analyzeforresumeuploadandpipelineinvocation;GET/api/applications/{id}/questionsforidempotentAI question retrieval;POST/api/behavioural/analyzefor NLPresponse evaluation;POST/api/applications/{id}/submitfor compositescoreandreportgeneration;GET/api/reports/{id}/pdfforbinaryPDFdelivery.

4.4

Frontend Implementation

The candidate-facing five-step application wizard manages state with React local state and TanStack Query mutations. React Dropzone powers the resume upload with MIME type filtering (PDF/DOCX). A visual score summary renders circularScoreRingcomponentswithcolour-codedthresholds:emeraldgreen(>70%),amber(40–70%),red(<40%).The recruiter analytics dashboard renders three Recharts visualisations: score distribution bar chart, application status pie chart,andtop-10candidateshorizontalbarchart.

5. RESULTS & DISCUSSION

5.1 Performance Metrics

Table-3presentsthesystemevaluationresultsacrosskeyoperationalmetrics.

Table-3: System Performance

Metric Value Notes

APIresponse (standard) <200ms

Indexed MongoDB queries

Resumepipeline(LLM warm) 8–15s CPU inference, Ollamaactive

Resumepipeline(LLM cold) 20–45s Model loading overhead

Resumepipeline(GPU) 2–5s

NVIDIA+ OllamaGPU backend

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Behaviouralanalysis 1–3s

PDFreportgeneration <1s

Lighthousescore >85/100

Semanticmatching (qual.) ~90%

JWTvalidation overhead <5ms

Unittestpassrate 16/16

Integrationtestpass rate 20/20

Functionaltestpass rate 15/15

5.2 Qualitative Analysis

VADER+ scikit-learn only

ReportLab synchronous

Code splitting,lazy loading

50test resumes, τ=0.70

In-memory HMACverify

Allmodules

FullAPI flowsend-toend

Bothuser role workflows

The semantic skill matching approach demonstrated clear advantages over keyword-based matching. Candidates describing 'REST microservices' were correctly matched against 'API development' at cosine similarity 0.78 a false negativeunderexact-keywordsystems.Similarly,'deeplearningmodeldevelopment'matched'neuralnetworkexpertise' at0.82similarity.

Thebehaviouralscoringmoduledemonstratedconsistentdifferentiationbetweenresponsequalitylevels.Responseswith high vocabulary diversity, positive VADER compound scores (>0.3), and adequate length (>50 words) received scores in the0.65–0.85range.Minimalorcopy-pastedresponseswereappropriatelycappedat0.12.

The composite scoring model, on informal review by domain experts, placed the most qualified candidates in the top quartileinthemajorityoftestcases.Theequal50/50defaultweightingproducedbalancedrankingssuitableforgeneralpurposehiring.

5.3

Discussion

NextHire demonstrates that integrating open-source AI components into a cohesive recruitment platform is technically feasibleandproducesmeaningfulimprovementsoverkeyword-basedscreening.Theprimaryperformancebottleneck CPU-basedLLMinferencelatencyof20–45soncoldstart isaddressablethroughGPUacceleration,reducinglatencyto 2–5s.ThetransparencyofNextHire'sscoringprovidesapracticalcomplianceadvantageasAIfairnessregulationsevolve globally.

6. CONCLUSIONS

This paper presented NextHire, a full-stack AI-powered recruitment platform addressing critical deficiencies of existing automatedscreeningsystems.Theplatformintegratesamulti-stageresumeanalysispipelineusingsentence-transformer

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072 © 2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008 Certified Journal | Page 1438

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

semantic embeddings and a locally hosted Qwen 2.5 7B LLM, with an NLP-driven behavioural assessment module employingVADERsentimentanalysisandscikit-learnpersonalityinference.

Evaluation results demonstrate sub-200 ms API response times, 8–15 s resume analysis on warm LLM instances, approximately 90% qualitative semantic matching accuracy, and a 100% pass rate across 51 unit, integration, and functionaltestcases.NextHiresuccessfullyaddressesthekeylimitationsidentifiedintheliterature:keyworddependency, siloedassessmentdimensions,opaquescoring,highdeploymentcost,cloudprivacyrisks,andminimalcandidatefeedback.

6.1 Future Scope

• GPU-acceleratedLLMinferencetoreduceanalysislatencyto2–5s.

• Multilingualsupportusingparaphrase-multilingual-MiniLM-L12-v2.

• WebRTC-basedasynchronousvideointerviewmodulewithtranscriptNLPanalysis.

• Fairnessmonitoringwithscoredistributionanalysisacrossdemographicgroups.

• ValidationofpersonalityscoringmodelagainsttheNEO-PI-3psychometricinstrument.

ACKNOWLEDGEMENT

The authors thank their project guide, the Department of Computer Science & Engineering, and the Apollo University, Chittoor,fortheirguidanceandsupportthroughoutthisproject.

REFERENCES

1. SocietyforHumanResourceManagement(SHRM),"TalentAcquisitionBenchmarkingReport,"SHRM,Alexandria, VA,USA,2022.

2. N. Reimers and I. Gurevych, "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks," in Proc. EMNLP,2019,pp.3982–3992.

3. T.Kuzman,I.Mozetic,andN.Ljubesic,"ChatGPT:BeginningofanEndofManualAnnotation?"arXiv:2303.03325, 2023.

4. P. L. Roth et al., "Social media in employee-selection-related decisions," J. Manage., vol. 42, no. 1, pp. 269–298, 2016.

5. T.Mikolovetal.,"DistributedRepresentationsofWordsandPhrases,"inProc.NeurIPS,2013,pp.3111–3119.

6. J.Pennington,R.Socher,andC.Manning,"GloVe:GlobalVectorsforWordRepresentation,"inProc.EMNLP,2014, pp.1532–1543.

7. J. Devlin et al., "BERT: Pre-training of Deep Bidirectional Transformers," in Proc. NAACL-HLT, 2019, pp. 4171–4186.

8. F. L. Schmidt and J. E. Hunter, "The Validity and Utility of Selection Methods in Personnel Psychology," PsychologicalBulletin,vol.124,no.2,pp.262–274,1998.

9. F.Mairesseetal.,"UsingLinguisticCuesfortheAutomaticRecognitionofPersonality,"J.Artif.Intell.Res.,vol.30, pp.457–500,2007.

10.C.J.HuttoandE.E.Gilbert, "VADER:AParsimonious Rule-basedModel for SentimentAnalysis," inProc.ICWSM, 2014.

11.M.A.McDanieletal.,"SituationalJudgmentTests,ResponseInstructions,andValidity,"Personnel Psychology,vol. 60,no.1,pp.63–91,2007.

© 2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008 Certified Journal | Page 1439

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

12.T.B.Brownetal.,"LanguageModelsareFew-ShotLearners,"inProc.NeurIPS,vol.33,2020,pp.1877–1901.

13.QwenTeam,AlibabaGroup,"Qwen2.5TechnicalReport,"arXiv:2412.15115,2024.

14.M.Raghavanetal.,"MitigatingBiasinAlgorithmicHiring,"inProc.ACMFAT*,2020,pp.469–481.

15.S.Bird,E.Klein,andE.Loper,NaturalLanguageProcessingwithPython.O'ReillyMedia,2009.

© 2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008 Certified Journal | Page 1440

Turn static files into dynamic content formats.

Create a flipbook
Next Hire: An AI-Powered Intelligent Recruitment Platform by IRJET Journal - Issuu