# Singapore's AutoTrust Posts Top Marks on Public Tests of Recursive Self-Improving AI With a Sliver of Rivals' Funding

- Link: https://www.thailand-business-news.com/pr-news/singapores-autotrust-posts-top-marks-on-public-tests-of-recursive-self-improving-ai-with-a-sliver-of-rivals-funding
- Published: 2026-09-28T22:00:00+07:00
- Author: PR Newswire

|   |

_Its ScienceGuru system led two public benchmarks and posted the fastest reported
times on two others in September_

SINGAPORE, Sept. 28, 2026 /PRNewswire/ — A startup with about 10 researchers spent
September setting the pace on public tests of AI that improves AI, a goal rivals
have raised hundreds of millions of dollars to pursue.

[⌊ScienceGuru's four benchmark results, September 2026⌉⌊ScienceGuru's four benchmark
results, September 2026⌉[
ScienceGuru’s four benchmark results, September 2026

AutoTrust AI said Monday that ScienceGuru, its agentic research system, posted results
on four public benchmarks used to test self-improving AI: first place on Autoresearch@
Home, the top validated entry on MedARC’s NanoPath when posted, and the fastest 
times reported to date on two GPT-2 training speedruns, both self-reported. Running
on AutoTrust’s own Guru Turbo models, ScienceGuru proposed, implemented and tested
each recipe; the company’s researchers set the objectives and supplied the computing
power.

Recursive self-improvement, or RSI—AI that improves how AI itself is built—is now
an explicit goal at the industry’s frontier. OpenAI said this month that, by its
own measurements, it had reached a goal set publicly last October: an automated 
AI research intern. The four benchmarks ScienceGuru tackled serve as public testbeds
for RSI.

"Recursive self-improvement is the race that will decide who builds the next generation
of AI, and it can now be measured in public," said Daniel Tang, AutoTrust’s co-founder
and chief executive. "In September, a small team in Singapore took first place on
two of those tests and posted the fastest reported times on the other two, with 
a fraction of the capital of the labs we are measured against."

**Four tests, one month**

The four results, in the order posted:

 ◦ **Autoresearch@Home (Sept. 1): First place on the official leaderboard. **Built
   on autoresearch, a project from OpenAI founding member Andrej Karpathy, the community
   benchmark gives an AI agent one GPU and five minutes of training to improve a
   language model. ScienceGuru’s entry scored 0.8895 validation bits per byte, where
   lower is better, and 20 reruns of the same code with a fixed seed averaged 0.8897.
 ◦ **MedARC NanoPath v2 (Sept. 9): Top validated entry when posted. **The benchmark’s
   maintainer independently retrained ScienceGuru’s pathology-AI recipe with three
   fresh seeds and validated it at 0.6597, where higher is better. When AutoTrust
   posted the result, the official leaderboard listed it as the validated leader(
   snapshot at 2:17 a.m. UTC on Sept. 9). It ranked second among validated entries
   as of Sept. 26.
 ◦ **NanoGPT Speedrun (Sept. 25): 24.90 seconds, 2.7 times as fast as the official
   record. **ScienceGuru trained GPT-2 Small to the 3.28 validation-loss target 
   of Keller Jordan’s community speedrun on AutoTrust’s own node of eight NVIDIA
   H100 GPUs, averaging 24.90 seconds over five fixed seeds with every run reported.
   That compares with the official record of 67.56 seconds. Self-reported; not yet
   reviewed by the benchmark’s maintainers.
 ◦ **Time-to-GPT-2 (Sept. 25): 72.24 minutes, 27% less time than the official record.**
   ScienceGuru trained a GPT-2-grade language model on a single 8×H100 node in 72.24
   minutes, against an official record of about 99 minutes on Mr. Karpathy’s nanochat
   leaderboard. That beats the best community recipe’s six-run average by 9.6 minutes
   and the fastest community experiment’s three-run average by 1.7 minutes. By AutoTrust’s
   review of public submissions, it is the fastest time reported to date; it comes
   from a single, self-reported run.

**How ScienceGuru did it**

ScienceGuru ran the same research loop on every benchmark: survey the published 
work and open submissions, combine the strongest ideas, implement and test them 
on the benchmark’s standard hardware, verify results with pinned source code and
complete evaluations, then publish code, verification materials and credit. All 
of it is at [github.com/AutoTrustAI](https://github.com/AutoTrustAI).

On the NanoGPT Speedrun, ScienceGuru fused two pending community submissions—ANVIL2
by Deven Pietrzak and Exact-match by Herman Brunborg—cut the training schedule from
1,194 steps to 652 and re-engineered how the server’s processors feed its GPUs. 
On Time-to-GPT-2, it started from nanochat and a community recipe by Giovanni Zinzi,
narrowed the model’s feed-forward layers to about 14% below nanochat’s default width
and set a 9,841-step training horizon. On Autoresearch@Home, it proposed and implemented
architecture, optimizer, kernel and memory-layout changes within the fixed five-
minute budget.

[⌊Time-to-GPT-2: reported times on Andrej Karpathy’s nanochat leaderboard and public
submissions⌉⌊Time-to-GPT-2: reported times on Andrej Karpathy’s nanochat leaderboard
and public submissions⌉[
Time-to-GPT-2: reported times on Andrej Karpathy’s nanochat
leaderboard and public submissions

**A fraction of the capital**

AutoTrust produced the results as a seed-stage lab that has raised less than $10
million to date. Other startups pursuing automated research have raised far more:
besides Recursive’s $650 million, Periodic Labs launched in 2025 with a $300 million
seed round. Of those labs, Recursive is the only one AutoTrust found with published
results on these benchmarks, and ScienceGuru’s figures are better on both tasks 
the two share: a self-reported 24.90 seconds on the NanoGPT Speedrun against the
75.36-second June record credited to Recursive’s system, and a leaderboard score
of 0.8895 on the five-minute autoresearch task against the 0.9109 average Recursive
published from its own runs.

**From benchmarks to business**

The benchmarks are a proving ground; the product is the system behind them. ScienceGuru
is the top layer of what AutoTrust calls an agentic operating system for research
and model development.

[⌊The AutoTrust stack: an agentic OS for research and model building⌉⌊The AutoTrust
stack: an agentic OS for research and model building⌉[
The AutoTrust stack: an agentic
OS for research and model building

The stack integrates four capabilities:

 ◦ **Proprietary models: **the Guru family of foundation models. Guru Turbo 1.0 
   and 1.2 powered all four September results.
 ◦ **Agentic harness: **ScienceGuru, which reads the literature and prior work, 
   forms hypotheses, writes and runs code, audits results and writes them up across
   long research sessions.
 ◦ **Model harness: **tooling that post-trains, evaluates, compresses and deploys
   models, including Blocks of Experts, LoRA, reinforcement learning and mixture-
   of-experts rewiring, plus Neural Architecture Searching, i.e pruning and quantization
   tuned to the target hardware.
 ◦ **RSI capabilities: **ScienceGuru supervises the training of Guru models alongside
   AutoTrust researchers, and verified research trajectories become training signal
   for the next generation of models.

That last layer is what makes the system recursive: ScienceGuru runs on Guru models
and helps train the next ones.

"What’s unique is the combination: our own models, an agentic harness that does 
real research, and model and memory harnesses that turn what it learns into better
models," Mr. Tang said.

ScienceGuru is available now for macOS and Windows at [scienceguru.ai](https://scienceguru.ai/),
and AutoTrust uses the same stack to build customized, sovereign models that enterprises
train on their own data.

"Every recipe behind these results was proposed, implemented and tested by ScienceGuru,
and the code is public so anyone can check it," said Josh Liu, AutoTrust’s co-founder
and chairman. "Customers get the same loop today: ScienceGuru for their research,
and the same harness to train and deploy models on their own data."

**About these results**

Standings are as of the dates shown and can change as new entries arrive. The NanoGPT
Speedrun and Time-to-GPT-2 results are self-reported by AutoTrust and have not yet
been reviewed by the benchmarks’ maintainers. Comparisons use figures published 
by each source, checked Sept. 24–26, 2026. The benchmarks are independent of AutoTrust.

**About AutoTrust AI**

AutoTrust AI Pte. Ltd. is an applied AI research lab headquartered in Singapore,
with an office in Silicon Valley. It builds the Guru family of foundation models
and ScienceGuru, an agentic platform for scientific research and model development,
and trains customized sovereign models for enterprises. Learn more at [autotrust.ai](https://autotrust.ai/)
and [scienceguru.ai](https://scienceguru.ai/).

---

  |  This article was produced by Cision PR Newswire, our trusted news partner. The views expressed and the content presented here are solely those of the author and may not fully reflect the opinions of Thailand Business News. |

---

 
**Read the original article :** [Singapore's AutoTrust Posts Top Marks on Public Tests of Recursive Self-Improving AI With a Sliver of Rivals' Funding ](http://www.prnasia.com/story/archive/5057805_AE57805_0?rand=184833)
