THE SACRED COMPUTER - СВЕЩЕНИЯТ СМЕТАЧ, a.k.a. ARTIFICIAL MIND: Thinking Machines, Creativity and Human Development. A pioneering virtual Research Institute for Artificial General Intelligence (AGI), Cosmism and Transhumanism, AI, Software, Interdisciplinary General Research, Creativity, Versatility, Universal Men ~ Twenkids etc. Created in 2000. (...)
На SIGI-2026 публикувах поредната версия на "Първата стратегия" - допълнено издание (от 256 до 283 стр. като подбирам още материал за следващата версия.
Пускам го като нов файл: Purvata_Strategiya_UIR_AGI_2003_Arnaudov_SIGI-2026.pdf
Реших да направя това издание след като се сетих за един въпрос, бележка под линия, за цената на "Витоша" (1961-1963 - през 1962 се извършва деноминация), и най-вече като препрочитах основния том на "Пророците" и писмата ми до Христо Крушков и др., докато работя и по неговото ново коригирано издание, а по-късно може би и допълнено през 2026 г.
Осъзнах, че множество писма и други статии и документи от началото и средата на 2000-те, вече публикуани в Основни том, в "Кратка хронология", "Подробна хронология ..." имат място и вна по-"прегледно" и "видно" място в "Първата стратегия", защото подсилват по още по-"графичен" начин вече дадените доказателства. В Основния том някои са във въведенията, но други са след 1600-та страница в подробната хронология.
Преписвайте тази историйца и пазете я да не изчезне!
Github:
Разширено издание от 28.6.2026 г. - 283 с.
Много допълнения спрямо предната версия от 9.6.2026 г.* Вече 283 стр. (от 256, +27). Много добавени писма от Основния том на Пророците: "5. Продължения на първата стратегия с конкретни научни и приложни насоки, публикувани между 12.2004 – 2.2008 г.: Вселена и Разум 5; Как двама всестранни младежи предвидиха пораждащия изкуствен интелект и принципа на работа на преобразителите между 2002 и 2005 г. – из писма между Илиян Георгиев и Тодор от януари-февруари 2005 г. Писма до Христо Крушков с повече обяснения и подробности от 11.2007 г.; писмо до Атанас Чанев от 12.2007; „Smarty – най-интелигентният речник в света“, 5.2007. Статията с дейностите и плановете ми от блог Изкуствен Разум от 2.2008 г.: „Творчески планове - какво правя, искам да правя, мисля си че правя... Частици от тях...“ . Коментар от 8.2008 г. за “Автоматична програмираща интелигентност“. Виж и писма до Людмила от 12.2007 г. в послеписа."
* В сравненията с Хасабис - включен още един откъс от писмото до А.Чанев.
* Бележка за деноминацията от 1962 г. относно бюджета на "Витоша" от 1.5 - 3 млн. (но неуточнено преди или след деноминацията 1:10).
Повече от статията "Творчески планове" ... добавяне на бел. че книгата е и на SIGI-2026.
* В изданието от 28.6.2026 включвам повече текст от началото на статията, където споменавам изрично „интелигентните асистенти“.
* Изследователската група в Уулвърхамптън, основана от Руслан Митков през 1995 г. http://clg.wlv.ac.uk/ (вече не е достъпен):
* Как работи разумът? Йерархичен самоорганизиращ се предсказател на бъдещето – научно представление .. - повече информация за събитието в ТУ София 2009 г.
(...)
* Версията от 9.6.2026 г. беше качена на SIGI-2025 и twenkid.com/agi
Yann Lecun cites a post, which is acknowledging that his ideas were correct etc.
A part of the concluding punch lines:
* It is "dime a dozen", but people decades older than me who *literally* repeated and ripped-off my suggestions, [strategy, plans, principles, theories, directions, conclusions, thoughts ...] and observations decades later, got prized with billions to *waste* and I am not even mentioned. They did it even in my own country, where one Bulgarian-Canadian became an "architect" of an institute in Sofia, with statements which were a *20 years late rip-off* of the above-cited essay, which were sold as "innovative" and ground-breaking :))), "for the first time in Eastern Europe" etc.
* As of "Dime a dozen"--> yes, or even "Five a dozen" --> The current "supercomputer" of my lab is called "PETAK I", where"Pet" means "5": from: 1. "Pentium" (historically the CPU and brand on which TUM was created), 2. The CPUs of all nodes: Core i5 (all old ones, 11-14 years old models, LOL); 3. Five nodes of the cluster (the initial full configuration) 4. A parody CPU-name from a science fiction work from 2004 from that theory ("Pentium 5"; "Петият Петак") and 5. In Bulgarian it also means "5 cents"... LMAO
Yann LeCun may have been right about something important: next-token and next-pixel prediction are probably not the most efficient path to real world understanding.
For years, the industry has been scaling generative models under the assumption that bigger models, more data, and more compute would eventually produce deeper intelligence. LeCun has been arguing the opposite: predicting every word or every pixel forces models to spend huge amounts of compute on surface details instead of learning the underlying structure of reality.
That’s the core idea behind JEPA (Joint-Embedding Predictive Architecture): instead of reconstructing the world pixel by pixel, learn a compact latent representation and predict what happens next inside that space.
The problem is that these models have historically been unstable. They suffer from “representation collapse,” where the latent space becomes too simple to carry useful information unless you add complex training tricks, auxiliary losses, or frozen components.
A new paper, LeWorldModel (LeWM), shows a much cleaner approach. It trains end-to-end from raw pixels using only two losses: a next-embedding prediction loss and a Gaussian regularizer on the latent space. This drastically simplifies the training setup compared to prior approaches.
The efficiency gains are striking. The model has around 15 million parameters, trains on a single GPU in a few hours, and can plan up to 48× faster than larger foundation-model-based world models, while staying competitive on several 2D and 3D control tasks. Its latent space also appears to capture meaningful physical structure and can detect physically implausible events in controlled environments.
This doesn’t mean generative AI is a dead end. LLMs remain extremely powerful. But it does reinforce a key technical point: for world modeling and physical reasoning, predictive latent-space approaches may be far more compute-efficient than brute-force generation.
The real shift might be this: not models that generate everything, but models that understand enough of the world to predict what actually matters."""
From the books "The First modern AI Strategy ..."... and "Stack Theory is Yet another fork of Theory of Universe and Mind" published at SIGI-2025
This idea, together with the prediction and next-token prediction (but in multi-scale, multi-precision hierarchy of resolutions of causality-control and perception), was published and explained nearly 25 years ago in Theory of Universe and Mind and presented during the world's first university courses in AGI in 2010 and 2011. Y.Bengio also rediscovered it 2017-2018 (Consciousness prior) and his example is almost literary repetition of an introductory definition from a treatise published about 14 years earlier. The author was a teenager, LOL.
Yann LeCun:
@Todor Arnaudov as I pointed out on another platform, ideas are a dime a dozen. The hard part, for something like this, is to implement it and to make it work.
The whole idea of hierarchical representations and learning by prediction is very old.
But learning hierarchies of representations didn't really work until convolutional nets were shown to do it in the late 1980s and more forcefully in the early 2010s (this took a while).
=== Todor Arnaudov:
Hi, first thanks for your answer as I didn't expect this honor. I don't disagree that there were earlier "prophets", I recently published a hyperbook with a related name (nearly 5000 pages in total), where one of the intros in one of the sections with collectons of related, prior and later work is a citation from the Holy Bible: "There is nothing new under the Sun"
Some of the prior work doesn't get enough credit and is unknown, even the "fellow AI historian" Schmidhuber doesn't mention them, e.g. the Soviet lab of Bongard and his colleagues etc. (E.g. once I caught Chollet literary restating insights from the 1967 book "Проблема узнавания" - perhaps he didn't know; he also rediscovers definitions for general intelligence of mine, published in 2001 (he couldn't know about it) - see the link at the end and the reviews of the LLMs).
The Bible is called "The Prophets of the Thinking Machines: Artificial General Intelligence & Transhumanism: History, Theory and Pioneers; Past, Present and Future", SIGI-2025 - and yes, almost nobody will bother to even open it. :))
BTW, e.g. IMO your PhD student Marc’Aurelio Ranzato deserves more credit for his pioneering work in DL and his insights (which perhaps [are] ~ also yours) -- his work is credited in my historical collections here: https://twenkid.com/agi/Lazar_The_Prophets_of_the_Thinking_Machines_20-8-2025.pdf ~p.21.
I do agree that I had to push to implementations immediately (not your type of NNs though) and perhaps my claims would be accepted after I implement them all by myself (Or if I or somebody else had - 20 years ago with no collaborators or any funding, no mechanical Turks to labe a gazillion of data and computing iterations, compared to 20 years later and all the collected resources in all senses of the word: i.e. IMO the difficulty of the implementation is supposed to decrease and be "discounted" with time like in RL; an idea 25 or 50 years ago may end up more "valuable" than an implementation in the present - see generative AI and the final citation below)
* I know about your dismissive opinion about "ideas", e.g. your comments to Schmidhuber's recent challenge, that you also could find ideas in your unpublished notes or something etc. and I've listened to your answers to him since 2022, "The path towards autonomous AI..." - I remember you defended yourself with referring to Optimal Control etc.
However many works are proposals, theoretical etc. but still get recognized, while other prior ones - don't and are even "humiliated". Also the core novelty there in my reading of the paper was also matching the mentioned TUM (and too general, it was not an implementation too); in general it looked like another cognitive architecture, which were popular in the cognitive science and the AGI community decades earlier, perhaps I have to reread it.
* I understand that if you dismiss even the German, who is at a comparable status as yours or, say he has more ground to be believed that he is, then you (and almost anyone) wouldn't recognize the claimed "priority" or even just the "contribution" of some obscure "self-proclaimed" "crank" or the mentioned theory, no matter the evidence (maybe you wouldn't even bother to check any evidence or count it as "theory" or anything).
BTW, your recent work about the brain/humans as "not general ..." also matches and is closely related to my prior work/accounts, beginning in early 2000s, however with different interpretation of the observations. The limitations don't deny the concept of general intelligence and the possibility of general principles and modules (prediction-compression etc.) I may address the correspondences in a paper.
* Stack Theory is yet another Fork of Theory of Universe and Mind, SIGI-2025
* The first modern AI strategy was published by an 18-year old in 2003 and repeated and implemented by the whole world 15-20 years later: Bulgarian Prophecies: How would I invest one million for the greatest benefit for the development of my country?https://twenkid.com/agi/Purvata_Strategiya_UIR_AGI_2003_Arnaudov_SIGI-2025_31-3-2025.pdf (Bongard, 1967 vs Chollet,2024 p.169-170)
* BTW, cheers from Kyuchuk Paris- that's the district in the city of Plovdiv, where TUM was created. 🙂
* This is the world's first modern "AI strategy", 2003, repeated and implemented by "the whole world" 15-20 years later: https://twenkid.com/agi/proekt.htm
* It is "dime a dozen", but people decades older than me who *literaly* repeated and ripped-off my suggestions and observations decades later, got prized with billions to *waste* and I am not even mentioned. They did it even in my own country, where one Bulgarian-Canadian became an "architect" of an institute in Sofia, with statements which were a *20 years late rip-off* of the above-cited essay, which were sold as "innovative" and ground-breaking :))), "for the first time in Eastern Europe" etc.
* As of "Dime a dozen"--> yes, or even "Five a dozen" --> The current "supercomputer" of my lab is called "PETAK I", where "Pet" means "5": from: 1. "Pentium" (historically the CPU and brand on which TUM was created), 2. The CPUs of all nodes: Core i5 (all old ones, 11-14 years old models, LOL); 3. Five nodes of the cluster (the initial full configuration) 4. A parody CPU-name from a science fiction work from 2004 from that theory ("Pentium 5") and 5. In Bulgarian it also means "5 cents"... LMAO
Also as I predicted in 2013 (counterintuitive to all "experts" up to just a few years ago, I namely wrote this article *because* of clueless "experts" predicted the opposite; they were later cited thousands of times for their *WRONG* world-model and wrong predictions):
"Creative Intelligence will be First Surpassed and Blown Away by the Thinking Machines, not the "low-skill" workers whose jobs require agile and quick physical motion and interactions with human-sized and human-shaped environment"
" (...) For the intellectual jobs - it's much easier to pick a computer, run the appropriate software or connect it to the service,
and get it thinking - you already have decent cameras, microphones and many sensors even in smartphones. (...) The bottom line is that the "white collars" are more endangered in current-time economy. Perhaps that kind of economy could hardly survive the AGI revolution. I guess it may turn upside down for a while - the low-skill workers could get higher pay, because intellectual activities will be done in 1 ms for free... 😉 We, the smart guys (the smart asses, see "Super Smartasses" the graphical series ) wouldn't be needed by anyone... Not that we are needed now. :))"
* The prediction of the generative AI (however it could have been created by the late 2000s-early 2010s - it came *too late*, not too quick as Hinton and Bengio "complain"; not with gradient-descent of course):
-- Creativity is Imitation at the Level of Algorithms - An outline sketch of a possible path of development of the Artificial Intelligence "Emil"
Updates to the table about the biggest GPT-models circa 2021 in the book "The First Modern Strategy ... "
The GPT2-MEDIUM-BG seems to be among the biggest 6-7 models, trained on a free single Tesla T4 in Colab. :))
p.26
В
този труд са дописани допълнителни бележки към цитирани откъси от класическия
ТРИВ от [17]
за
мерките за зародиш на разум и степените на развитие. От[TT1] големите езикови модели –
сведения за тях и работата им и някои най-нови публикации, както и сравнение на
данни за ранни GPT-модели на различни езици – арабски, френски и множество европейски,
японски и китайски – като българският GPT2-MEDIUM-BG се оказва един от шест-седем най-големи модели
от такъв тип в света до 2021 г. за езици различни от английския – по-големи
около същото време или малко по-рано са само за китайски, арабски, руски, румънски
и френски; с подобен размер е за японски[1],
разработен по същото време като българския. Кратки бележки за по-големия проект за
инфраструктура за Общ ИИ и всякакви проекти, свързани и с пораждащи модели – „Вседържец[TT2]“.(...)
[1] Възможно е и др.: на 18.5.2026 добавих
китайски, руски и нов испански от 2022 г. Виж [236]
p.238 [236]
236. Ранни
пораждащи големи езикови модели от типа GPT за езици, различни от английския:български, френски, арабски, испански, португалски, немски, китайски; гръцки, сръбски, румънски, японски, китайски,
руски – 2020-2021 г. Датата на някои –
по дати на файловете с теглата на модела, дата на научна статия и пр. До края
на 2021 г.само китайският, френският, арабският, руският,
румънският, японският и българският са с над 100-тина милиона параметъра.
Румънският е силен, обучаван на 17 GB-ов корпус. Само българският
вероятно е разработен от един-единствен човек с бюджет и подкрепа = 0 и авторът
представя родната компютърна лингвистика в тази дисциплина като самозван
„хайдутин“, понеже институциите и по-„елитните“ бойци чакаха до 2023-2024
г. [66]. Сравни с аналогичен случай с ДЗБЕ
около 2001-2003 г. и бездействието на ИБЕ на БАН и на останалите филолози от
университетите спрямо явленията, срещу които ДЗБЕ се противопоставяше и се
опитваше да „призове“ „чети“ [16][40], а „маститите“
езиковеди (по определението на Павлин Стойчев, „PC World Bulgaria“, 5.2003[239]) гледаха безучастно и обясняваха, че това
били „естествени процеси“. Сравни с бележките за „Добродетелната
дружина и нехранимайковците“и [40], 2003 г., дали талантите не са имали избор да не
учат в „най-престижните университети“ и да развият местните и пр. XLM-R от „Фейсбук“, 11.2019 е по-голям, но в него
българският е един от 100 езика, на които е обучаван, и моделът е за
класификация и отговаряне на въпроси, а не за пораждане. Таблици:подредени по време на създаване и по
размер: Допълнена на 18.5.2026 с китайския,
руския и испанския голям модел.
Ранни големи
езикови модели “GPT“за разни езиципо време
15. Unsupervised Cross-lingual Representation Learning at Scale, Alexis
Conneau, Kartikay Khandelwal, …, Veselin Stoyanov, 11.2019/4.2020 - XLM-R многоезичен
езиков модел, обучаван и върху български корпус. В разработката участва Веселин
Стоянов.
[Updates in the edition from 17.5.2026: Chinese, Russian, Spanish] 16. CPM: A Large-scale Generative
Chinese Pre-trained Language Model, Zhengyan Zhang, Xu Han, Hao Zhou, Pei Ke,
Yuxian Gu, Deming Ye, Yujia Qin, Yusheng Su, Haozhe Ji, Jian Guan, Fanchao Qi,
Xiaozhi Wang, Yanan Zheng, Guoyang Zeng, Huanqi Cao, Shengqi Chen, Daixuan Li,
Zhenbo Sun, Zhiyuan Liu, Minlie Huang, Wentao Han, Jie Tang, Juanzi Li, Xiaoyan
Zhu, Maosong Sun, 1.12.2020, https://arxiv.org/abs/2012.00413
17. Methods for Detoxification of Texts for the Russian Language Daryna
Dementieva‡ , Daniil Moskovskiy‡ , Varvara Logacheva‡ , David Dale‡ , Olga
Kozlova† , Nikita Semenov† , and Alexander Panchenko‡ ‡Skolkovo Institute of
Science and Technology, Moscow, Russia †Mobile TeleSystems (MTS), Moscow,
Russia {daryna.dementieva, daniil.moskovskiy, v.logacheva, d.dale,
a.panchenko}@skoltech.ru {oskozlo9,nikita.semenov}@mts.ru, 19.5.2021 https://arxiv.org/pdf/2105.09052https://github.com/ai-forever/ru-gpts
18. Spanish Language Models, Asier Gutiérrez-Fandiño, Jordi
Armengol-Estapé, Marc Pàmies, Joan Llop-Palao, Joaquín Silveira-Ocampo,
Casimiro Pio Carrino, Aitor Gonzalez-Agirre, Carme Armentano-Oller, Carlos
Rodriguez-Penagos, Marta Villegas; 15.7.2021 – 5.4.2022 (v1 to v5); the GPT
models appears in v3 from 1.4.2022 https://arxiv.org/abs/2107.07253v3
* GPT2 е обявен от OpenAI през
2.2019 г., но не е публикуван за използване заради опасения за възможна
злоупотреба – пораждане на „фалшиви новини“ и пр. През 8.2019 пускат 774М, а
през 11.2019 – двойно по-големият. OpenGPT2,
представен през 8.2019 г., е обучен върху корпуса „OpenWebText”;
цената за облачни услуги била около 50 хил. долара. https://en.wikipedia.org/wiki/GPT-2
* Размерите са приблизителни и може да са неточни за модели, които не са
ползвали точно архитектурата, някои са с различен брой токени (32000) и пр. Влияние
оказва не само броят параметри, а още качеството на данните и начинът на
обучение и др. На онзи етап и мащаби всички модели са експериментални и с
научна и образователна цел.
* [Submitted on 1 Dec 2020]
CPM: A Large-scale Generative Chinese Pre-trained Language Model
during training: batch
size = 3,072; 3M tokens .. (vs 1 M for GPT3 training) strong few shot learning ...
* Methods for Detoxification of Texts for the Russian Language
Daryna Dementieva‡
, Daniil Moskovskiy‡
, Varvara Logacheva‡
, David Dale‡
,
Olga Kozlova†
, Nikita Semenov†
, and Alexander Panchenko‡
‡Skolkovo Institute of Science and Technology, Moscow, Russia
†Mobile TeleSystems (MTS), Moscow, Russia
{daryna.dementieva, daniil.moskovskiy, v.logacheva, d.dale, a.panchenko}@skoltech.ru
{oskozlo9,nikita.semenov}@mts.ru, 19.5.2021
https://arxiv.org/pdf/2105.09052
https://github.com/ai-forever/ru-gpts
* Others: later than GPT2-MEDIUM-BG (2021)
Spanish - MarIA GPT-2 -
https://arxiv.org/pdf/2107.07253v1 - only BERT-like model in July 2021
GPT2 models, up to Large 774M appears in v3, 1.4.2022:
https://arxiv.org/abs/2107.07253v3
...
German - https://www.kkirchheim.de/blog/german-gpt/ - "Training a German LLM from scratch"
Existing German models available on Hugging Face have 137M parameters and a context length of 1024 tokens1, which is quite limited compared to recently released models, such as those in the LLAMA family.