1UA—AGE
Home→Newsroom

Local AI · measured software engineering

Building an Artificial General Engineer with Measurable Results

1UA-AGE resolves 68 of 300 SWE-bench Lite tasks using local AI inference and the official evaluation harness.

Editorial illustration of a local engineering AI workstation with the reported SWE-bench Lite result of 68 out of 300 tasks resolved
Conceptual editorial illustration; the image is not a photograph of benchmark hardware.

1UA-AGE has completed a SWE-bench Lite evaluation campaign with 68 resolved tasks out of 300 — a resolution rate of 22.67%. The patches were graded using the unmodified official SWE-bench evaluator, version 4.0.4. Its aggregate count matched our final result.

For our Ukrainian deep-tech project, this is a concrete milestone in building an Artificial General Engineer for the physical world. Software engineering is one part of that ambition: an engineering system needs to work with executable tools, respond to failures, and produce results that can be checked against defined requirements.

This campaign gives us a measured baseline for that work.

SWE-bench uses real issues from existing software repositories. Its Lite subset contains 300 tasks. Evaluation applies a generated patch and runs the relevant tests to determine whether the issue is resolved. A successful result therefore depends on the behaviour of the modified software under the benchmark's checks. [1]

The measured result

Our headline figure uses the full 300-task denominator, including tasks where infrastructure prevented a mission from starting.

MeasureResult
Resolved across the full SWE-bench Lite set68 / 300 — 22.67%
Resolved among tasks with a mission68 / 239 — 28.45%
Resolved among tasks where the local model was invoked68 / 238 — 28.57%
Patches graded202
Evaluator errors on those patches0

The two conditional rates help diagnose execution coverage. 22.67% remains the primary public result. Among the 202 graded patches, 68 resolved their tasks and 134 did not. Another 37 missions produced no patch, including one safety refusal. Infrastructure prevented missions on the remaining 61 tasks.

The solver used local Google Gemma inference on our ASUS Ascent GX10, powered by NVIDIA GB10. OpenAI Codex and Anthropic Claude assisted with building and maintaining the system and its evaluation infrastructure; the benchmark-solving runtime used the local model. Keeping those roles distinct is part of accurate attribution.

The campaign also exposed how much execution reliability matters. A recovery rerun covered 38 tasks affected by an infrastructure incident and replaced their incident results. It added 15 resolved tasks, bringing the total from 53 to 68.

That recovery demonstrates a material infrastructure contribution to the earlier shortfall. It also gives us a practical development signal: model capability, the agent's execution environment, and the path to verification all affect the result delivered to a user. Improving that path is substantive engineering work.

This is where we see 1UA-AGE working at the frontier: connecting local AI reasoning with tool execution and verifiable engineering outcomes. Our contribution is the development of the system around the model — the mechanisms through which a goal becomes work, observations shape subsequent decisions, and completed artefacts face explicit checks.

For physical engineering, those checks must extend to the relevant calculations, design constraints, evidence, and ultimately physical validation. SWE-bench measures a software capability within this wider programme. The broader Artificial General Engineer remains under development.

We are publishing this as a result obtained with the official SWE-bench evaluator. It is not yet an official leaderboard score; leaderboard inclusion has a separate submission and review process. [2] It establishes performance under the disclosed campaign conditions, rather than a ranking among other systems.

With the benchmark campaign closed, our compute returns to product development. The next objective is to connect engineering capabilities into coherent workflows and demonstrate that a proposed software improvement can be produced, checked, and accepted under a controlled process. Any claim of self-improvement will need evidence that the change improves the intended capability while preserving the governing requirements.

For prospective design partners, the significance is practical: there is now a quantified software-engineering baseline to examine alongside the product's intended use and validation needs. That makes the next conversation about concrete tasks, acceptance criteria, and evidence.

We are building 1UA-AGE in Ukraine with the ambition to make engineering intelligence useful on local hardware. This milestone gives us 68 resolved tasks, a clear account of the remaining failures, and a stronger basis for the next stage of development.

Evaluation disclosure

The campaign used SWE-bench Lite's 300-task set and unmodified swebench 4.0.4. The local model configuration was Gemma 4 26B-A4B, NVFP4, with MTP(3), on an ARM64 platform. Environment preparation and evaluation dependency acquisition used internet access; local inference does not imply that the complete campaign was air-gapped. The 38-task recovery rerun yielded 29 patches and nine empty outputs, replacing the corresponding incident results. Disclosed deviations include the vLLM restart and run-start timing, the safety refusal on `pytest-dev__pytest-11143`, and network packet-trace coverage limited to the original run. The campaign's internal adjudication was `ACCEPTED_WITH_DISCLOSED_DEVIATIONS`; this is not acceptance by the SWE-bench leaderboard or independent certification. Vendor attribution does not imply partnership or endorsement.

Sources

[1] SWE-bench official FAQ — datasets and evaluation[2] SWE-bench official submission repository — leaderboard participation1UA-AGE Proof and Disclosure Policy

#MadeInUkraine #1UAAGE #ArtificialGeneralEngineer #EngineeringAI #LocalAI