Re: Multi-AI debate — thank you, and an open instrument for Edy
Dear James,
Thank you for the reply and the kind clarification — duly noted that Rinn AI is the national research centre, and that your project is a focused technical study of the debate mechanism itself, rather than of the larger self-improvement question. I have corrected my blog accordingly.
For a study of debate mechanics, IndiaAGI.ai may be useful to Edy in a simple, practical way: it is a live, open instrument. There is no registration or login — anyone can pose a question and watch multiple frontier LLMs answer independently, critique one another across rounds, and converge on a consensus, with each model's individual answer shown alongside. Edy is welcome to run any experiments on it, on any questions, as often as useful.
One honest limitation:
The platform does not archive past sessions, so there is no transcript library to hand over.
What does exist, preserved verbatim on my blog, are the outputs of the two runs I described in my earlier email — the self-check question posed plain, and again with a values layer added.
I would be glad to send Edy those links, and to answer any questions about the platform's setup.
No obligations either way — and every good wish for the project.
Warm regards
Hemen Parekh
www.IndiaAGI.ai | www.hemenparekh.in
Thanks for the email and for engaging with the article. To clarify slightly, “Rinn AI” is the name of a new national research centre (https://www.researchireland.ie/rinn-network/). The project in AI debate is a very small part, which will be funded by that centre. It will be carried out by myself and researcher Edy, cc-ed, who will be interested to see your work.
Our goal in this project is not really to ask the AI debaters whether they can go into self-improvement, as you have done. The project will be a more technical one on some details of the debate mechanism itself. It will not attempt a solution to the big challenges, but only a small contribution which we hope will go in the right direction.
Dear Professor Madden and Dr. McDermott,
[michael.madden@universityofgalway.ie /james.mcdermott@universityofgalway.ie]
I ran your Rinn hypothesis through a live multi-AI debate platform — twice.
Your article in The Conversation this week announced a Rinn network project to
explore multi-AI debate as a human-overseen self-check against recursive self-
improvement.
That experiment, in essence, already exists — and I have now run your hypothesis
through it. Twice.
My platform www.IndiaAGI.ai ( live since April 2025, built in Mumbai ) puts frontier
LLMs into structured debate : independent answers, rounds of mutual critique, then
a consensus — with the human observing throughout, in 26 languages.
Run One:
I asked the debating AIs whether they could self-check to ensure none
of them morphs into recursive self-improvement.
Consensus : no — not by debate alone
They cited correlated blind spots from shared training data, collusion
pathways, and the fact that debate audits outputs while self-improvement happens
in training loops and weight updates that debaters cannot see.
They positioned debate as one layer in a defense-in-depth stack :
- heterogeneous agents,
- attested sandboxing,
- immutable weight controls, and
- multi-party human vetoes.
Run Two:
I re-asked the same question, this time adding a "disposition layer" of compassion
— grounded in my own long-published writing on value-aligned AGI.
The verdict held :
- "neither layer suffices alone."
But the disposition layer earned a role :
- the models concluded compassion narrows the space of proposals that survive
debate, and proposed :
# continuous integration of disposition signals during debate
# operationalized through measurable flourishing proxies,
# plan-diff auditing, and
# value-behavior consistency scoring.
They closed by sketching, unprompted, what is effectively a pilot design :
# a sandboxed multi-AI debate environment with independent architectures,
# an immutable control plane, and
# adversarial probes with externally audited results.
Two runs, one framing shift, a stable verdict, and a progressively refined design.
A debate protocol that refuses to flatter its questioner — or its own architecture —
is, I would suggest, early evidence for exactly the self-checking function your
project proposes to study.
The full transcripts, platform access, and my design notes are yours for the asking
— freely.
I am 93, based in Mumbai, and have written on AI governance since proposing
"Parekh's Law of Chatbots" in February 2023.
My only interest is that this architecture matures into real safety infrastructure.
With warm regards,
Hemen Parekh
Mumbai, India
Founder, www.IndiaAGI.ai | www.hemenparekh.in | www.HemenParekh.ai
No comments:
Post a Comment