makeitcount research — AI Software Maintenance: Scientific Study

makeitcount research — AI Software Maintenance: Scientific Study

Success is about trust. And trust is built through work done.

Geseidl Consulting Group

makeitcount research

At Geseidl, makeitcount isn't just a motto — it's a working principle. Every project, every line of code, every decision must count. When we saw that artificial intelligence could transform software maintenance, we didn't just use it — we documented it rigorously, measured it, and published it.

The result is a complete scientific paper: "Hybrid AI+Human Software Maintenance: A Case Study on mRemoteNG" — applied research demonstrating, with verifiable data, how an AI orchestrator can take over maintenance of an open-source project at industrial scale.

Why do we publish research? Because the best business decisions are made based on data, not intuition. We publish our results so anyone can verify, replicate, and improve upon them. Intelligent solutions for business development — built on measurable results.


The Scientific Paper

Title

Hybrid AI+Human Software Maintenance: A Case Study on mRemoteNG

Hypothesis

A supervised AI orchestrator can resolve over 50% of a legacy codebase's open issues, with a regression rate below 2% and a cost per commit below $5, outperforming manual-only development.

Subject

mRemoteNG — an open-source remote connections manager for Windows with 16 protocols (RDP, VNC, SSH, Telnet), ~114,000 lines of code, .NET Framework architecture with WinForms and COM references. 843 unresolved issues on GitHub at study start.

Methodology

The study went through 4 architectural generations of human-AI collaboration, over ~3 weeks, with a total cost of ~$320 in AI APIs. Instruments used:

  • Orchestrator: ~6,100 lines Python script coordinating AI agents
  • Supervisor: ~800 lines Python, handling 12 automatic failure modes
  • AI Agents: Codex Spark (OpenAI), Claude Sonnet/Opus (Anthropic), Gemini Pro (Google)
  • Independent verification: MSBuild + 6,175 automated tests (no AI)
  • Code quality: 5 levels — Roslynator, Meziantou, SonarCloud, CodeQL, Qodo AI Review

Results — Hypothesis Confirmed

All three hypothesis dimensions were confirmed:

83.3%Resolution rate702 of 843 issues (target: >50%)1.2%Regression rate7 of 585 implementations (target: <2%)$1.49Cost per commitStabilized from day 4 (target: <$5)

Cost per commit evolution

Day 1 $4.02Day 2 $3.01Day 3 $2.14Day 4+ $1.49

AI model performance

ModelSuccessMedian time% CostOptimal roleCodex Spark86%274s18%Primary executor — fast, cheap, reliableClaude Sonnet—474s58%Complex multi-file fixesClaude Opus—180s—Supervision, planning, strategic decisionsGemini Pro—417s24%Bulk transformations (rate-limited)

Code quality — complete transformation

Static analysis warnings
5,247 → 0 (100% eliminated)Automated tests
2,179 → 6,175 (0 failures)SonarCloud Quality Gate
PASSED — A / A / A


Four generations of learning

The path from "let AI fix everything" to a working system went through 4 architectural generations. Each failure generated new rules injected into agent prompts.

Gen 1 — Brute force

One AI agent, manual triage, 26 issues/sprint, 8h/day human attention. The baseline.

Gen 2 — Multi-agent orchestra

Three agents, fatal double-pay problem (27% of cost wasted on retries).

Gen 3 — The 31-hour disaster

201 phantom test runs accepted as passing. 353.7% pass rate unquestioned. BitDefender quarantined the DLL. Lesson: an unmonitored orchestrator is more dangerous than manual work.

Gen 4 — Self-healing supervisor

The architecture that delivered results. Feb 27 session: Codex alone resolved 89/104 issues (86%), zero human intervention.


Access the complete paper

📄 Hybrid AI+Human Software MaintenanceA Case Study on mRemoteNG — Geseidl IT Solutions, 2026PAPER.mdFull paper — hypothesis, methodology, results, discussionMETHODOLOGY.mdFormal methodology, instruments, metrics, baselineCOST_ANALYSIS.mdDetailed cost analysis by model and categoryFAILURE_CATALOG.mdPost-mortems: 31h disaster, 7 regressions, repo deletionEVIDENCE.mdVerifiable evidence trail (git, CI, SonarCloud)→ Read the paper on GitHub


Why this matters for business

This research isn't an academic exercise. It's proof that software maintenance can be economically transformed: ~$1.50 per commit, 24/7 operation, zero burnout. The model is reproducible — any project with issues, tests, and a build system can benefit.

For our clients, this means: faster legacy application modernization, predictable costs, quality verifiable through SonarCloud and VirusTotal. Not promises — metrics.

Discover how we apply this expertise in our IT services or read the complete mRemoteNG case study.

Need IT services and software modernization? The Geseidl Consulting Group team, CECCAR Prahova leader for 18 consecutive years, is ready to help. Discover our services or contact us for a free consultation.

Geseidl Consulting Group

CECCAR #1 Prahova · CAFR Rating A · ANEVAR · CCF #233 · ISO 9001:2015

Learn more about us →

Professional Accreditations

CECCAR #1 Prahova
CAFR Rating A
ANEVAR
CCF #233
ISO 9001:2015
ASPAAS
ANPC SAL - Solutionarea Alternativa a LitigiilorANPC SOL - Solutionarea Online a Litigiilor

Copyright ©2017-2026 Geseidl Consulting Group. All Rights Reserved.