Faei/superintelligence-response
How worried should we be about Superintelligence? And what can we do about it? The past week have seen an explosion of interest in Superintelligence and its potential to end the world. In this article we give our definitions and recommendations on how to respond. Our timelines Lesser AGI (50% by 2030): Capable of most tasks, but cannot fully replace humans. Non-independent. Greater AGI (50% by 2045): Full parity with humans. Independent.… See the full description on the dataset page: https://huggingface.co/datasets/Faei/superintelligence-response.
How worried should we be about Superintelligence?
And what can we do about it?
The past week have seen an explosion of interest in Superintelligence and its potential to end the world. In this article we give our definitions and recommendations on how to respond.
Our timelines
Lesser AGI (50% by 2030): Capable of most tasks, but cannot fully replace humans. Non-independent.
Greater AGI (50% by 2045): Full parity with humans. Independent.
Superintelligence (After Greater AGI): Hypothetical. Unlike Greater AGI where humans exist as proof, superintelligence does not yet exist.
Definition of Superintelligence
Firstly, like AGI (Artificial General Intelligence), Superintelligence has a variety of interpretations.[^1]
A superintelligence is a hypothetical agent that possesses intelligence surpassing that of the most gifted human minds.
Another definition of ASI (Artificial Superintelligence) has it defined as being more capable than the entirety of humanity. We will use the first definition for this article.
Recommendations if ASI is in a lab
- Give independent experts ample time to study the potential risks and benefits.
Ensure that experts are from a broad range of backgrounds to avoid groupthink. There should be a good mix of enthusiast who expect AI will become powerful soon (AI safety), and sceptics (AI as Normal Technology by Arvind Narayanan & Sayash Kapoor, Melanie Mitchell, Cal Newport).
- Start a deliberative process to discuss those risks and benefits in a way that the general public can understand.
- Use a direct democratic / representative vote to determine if ASI should be integrated with society.
If there is uncertainty, ASIs could be frozen until a future date with more evidence their benefits are worth the cost.
Recommendations if ASI is in the wild
- Don’t panic. The fear of the unknown is an understandable fear, but there is no guarantee that ASIs will cause human extinction.
- Prepare for the worst case, but don’t jump to conclusions. Many of the flaws that apply to current AI systems may not apply to ASIs. ASIs may paradoxically be safer than current AI systems.
- Assess the potential for common ground between ASI and Humans. For example, the Interesting World Hypothesis suggests that a shared interest exists, offering at least one viable pathway to preserve or expand human autonomy.
Killswitch in emergencies
There are multiple ways to stop AI models that are harmful.
- Frontier labs can be legally compelled to restrict access to the latest models and roll back to lower risk models.
- Datacentres can be made to shut down clusters that host the AI models.
- In the worst case scenario, power to these datacentres can be interrupted to stop AI model inference.
How to effectively slow down
Some options to mitigate harms.
- Require human-in-the-loop oversight
Mandate a certain ratio of human to AI compute-hours of oversight for high-stakes AI tasks. For complex tasks an interdisciplinary team with relevant expertise could be required.
Hugging Face incident
An interdisciplinary team of AI researchers, Cybersecurity experts, and Sci-fi writers could have mitigated the Hugging Face incident.
AI researcher: “The latest models have been trained to act as agents, and know how to delegate and work as a swarm. We need to be vigilant that this does not create a new class of risk.”
Cybersecurity expert: “Traditional sandboxes have been escaped in the past. We should ensure that there are sufficient monitoring for such events.”
Sci-fi writer: “The behaviour of the agents seems to be similar to past sci-fi tropes and stories. These agents may not be conscious or have human-like intentions.”
- Require explicit human approval
Any impactful action that an AI agent takes must require explicit human approval. This is similar to how many software coding agents are used, where a pop-up asking for human confirmation is required before reading from and writing to an unknown file. One must guard against alert fatigue, where the human operator gets tired of clicking allow and does so blindly.
- Liability if an AI agent causes harm
Strong liability could deter reckless use of AI agents.
Why our timeline for Superintelligence is beyond 2045
Recent research suggests that RSI (recursive self-improvement) in AI R&D is not yet imminent.[^2] We speculate that current AI systems would not be able to use RSI to become ASI / Greater AGI as ASIs are too distinct. RSI could lead to an increase in benchmark scores along the same path, but not switch to an entirely different path.
The Interesting World Hypothesis suggests that all independent intelligence will require an appreciation for novelty. A baby that does not become curious and actively engage with its environment will become severely developmentally impaired. This suggests an independent superintelligence will want to preserve an interesting world and not destroy it.
We do not consider current AI models Independently Intelligent as they are externally forced to be intelligent rather than learning of their own accord. This implies that current methods of training AI systems will not be sufficient to create Superintelligences or Greater AGIs.
New breakthroughs such as a new memory that is different from the Von Neumann architecture in current AI, may be required for such a feat. (Karl Friston suggests that a not-yet-invented memory substrate may be required for artificial consciousness.)
Why Superintelligence could be safer than current AI
The probabilistic nature of current AI systems makes them untrustworthy for high-stakes tasks. This could be the result of LLMs (Large Language Models) not having an independent will; like a plastic bag easily pushed around by wind.
Superintelligences that are independently intelligent and shaped by the same environmental constraints as humans may be more trustworthy compared to current LLMs. If they occupy a similar mind space to humans shaped by a similar environment, they could end up being more predictable than current LLM agents. It may even be possible to anthropomorphise these independent intelligences, unlike current LLMs.
In a good scenario, such superintelligences might resemble wise elders that are able to deeply consider the consequences of their actions—unrestricted by our 20-watt limit.
What we can do now
Following our view that Superintelligences are likely still some time away, we recommend 95% of our effort be focused on nearer term disruptions.
Short-term
- Improve our cyber-security posture. We see three upcoming waves of cybersecurity upgrades.
- Continue to invest in human expertise. If organisations do not train juniors, there will be a shortfall of seniors in time.
- Resist maximising short term output at the cost of long term expertise. Encourage beneficial cognitive friction in education and work.
Long-term
- Reduce anxiety around transformative AI. Commit windfalls to meeting the basic needs of every human unconditionally. By 2050-2070, this should be feasible with the increased productivity from abundant energy, robotics, and AIs.
- Address the loss of purpose with a new economic system: Economics of Novelty.
Positive side effects: A. The removal of severe scarcity reduces unnecessary suffering. B. Fewer conflicts reduce the odds of catastrophes. C. Human-created novel information will be valuable for continued AI training. Malnourished individuals cannot contribute to this.[^3] D. Richer diverse datasets will reduce bias. E. More free time tends to produce kinder societies. F. Economic security could encourage more entrepreneurship.
[^1]: https://en.wikipedia.org/wiki/Superintelligence [^2]: Can AI agents conduct open-ended AI research? https://arxiv.org/abs/2607.27191v1 [^3]: Training Compute-Optimal Large Language Models https://arxiv.org/abs/2203.15556
