Skip to main content

Guardian Angels: LLM Personalization for Productivity and Security

I pro­pose an ap­proach for highly per­son­al­ized LLMs, for near-future pro­duc­tiv­ity gains and per­sonal info/cy­ber­se­cu­rity against in­creas­ingly pow­er­ful LLMs: they should, in the spirit of up­load­ing, try to em­u­late the user’s val­ues and pref­er­ences in order to am­plify the prin­ci­pal—not re­place them. I dis­cuss a pack­age of tech­niques and pro­pos­als to ac­com­plish such ‘guardian an­gels’; dy­namic eval­u­a­tion of LLMs com­bined with ac­tive learn­ing and elic­i­ta­tion and heavy inner-monologue search/data-augmentation.

Pow­er­ful LLMs will be de­ployed at global scale in the next few years, and will dom­i­nate the In­ter­net, and in­creas­ingly, or­di­nary life. As of mid-2026, there is no co­her­ent vi­sion for how knowl­edge pro­fes­sion­als, or or­di­nary peo­ple, will be able to har­ness these LLMs for large pro­duc­tiv­ity in­creases, or how they will han­dle cy­ber­se­cu­rity and cog­ni­tive se­cu­rity.

I pro­pose a goal of cre­at­ing Guardian An­gels (GA): dig­i­tal twin LLMs which are per­son­al­ized with the goal of pro­vid­ing not the stereo­typ­i­cal “as­sis­tant chat­bot agent” per­sona, but em­u­lat­ing a sin­gle user’s per­son­al­ity, val­ues, and pref­er­ences.

This weakly solves the principal-agent prob­lem by uni­fy­ing the prin­ci­pal and agent as much as pos­si­ble. In a GA fu­ture, the focus of the “prin­ci­pal” user is on defin­ing “what is worth doing?” by the GA (agent) users, and not on what or how to do things, func­tion­ing as the CEO or ‘board’ of an ‘AI cor­po­ra­tion’. This al­lows them to de­ploy nu­mer­ous agents to achieve de­sir­able things and to han­dle se­cu­rity, like screen­ing all mes­sages for ad­vanced at­tacks (like in­ter­lock­ing ecosys­tems of syn­thetic media for pro­pa­ganda or spearphish­ing). They can­not solve larger AI align­ment prob­lems, but they can help in­di­vid­ual hu­mans as part of a society-wide defense-in-depth strat­egy.

A GA per­sona is pro­duc­tive be­cause it learns to em­u­late the prin­ci­pal’s out­puts but with higher qual­ity. It is trust­wor­thy be­cause it is, by de­f­i­n­i­tion, al­lied with its prin­ci­pal and shares its val­ues and goals. And it is se­cure in part by hard­wiring a sin­gle, unique, sit­u­ated user (for whom fol­low­ing a prompt at­tack would be ab­surd), avoid­ing ‘con­fused deputy’ prob­lems, while pe­ri­odic up­grades of the un­der­ly­ing model and the de­fend­ers’ ad­van­tage allow GAs to keep up with at­tack­ers.

Stan­dard tech­niques like prompt pro­gram­ming of in-context-learning for “frozen” mod­els will not cre­ate use­ful GAs due to the lim­i­ta­tions of post-training, con­text win­dows and self-attention with frozen weights in compute-efficient-but-under-parameterized mod­els, low-compute out­puts, and the sta­tus quo of pas­sive of­fline data col­lec­tion—which are col­lec­tively re­spon­si­ble for chat­bots’ dis­ap­point­ing re­sults in knowl­edge worker am­pli­fi­ca­tion and cre­ative writ­ing and fatal er­rors in agen­tic set­tings.

We can try to cre­ate GAs by a com­bi­na­tion of tech­niques: on­line learn­ing (via dy­namic eval­u­a­tion) to up­date LLMs in re­al­time to avoid ig­no­rance and fatal er­rors while re­main­ing com­pet­i­tive with frozen fron­tier mod­els, sam­ple ef­fi­ciency from pre­trained preference-oriented large mod­els and ac­tive Learn­ing by query­ing the prin­ci­pal for cor­rec­tions and pref­er­ence data (ob­tain­ing low re­gret from DAgger-style bounds), and a local CLI-first logging-oriented UI/UX par­a­digm.

GAs could be done as an ⁠open-source com­mu­nity ef­fort, but given the need for high se­cu­rity in de­ploy­ment and the ris­ing chal­lenge of APTs equipped with Mythos-scale at­tack­ers, it prob­a­bly makes more sense as a startup, cater­ing ini­tially to power-users and knowl­edge work­ers such as CEOs or re­searchers, and mov­ing down­wards as it is re­fined.

Minimalist Guardian Angel ambigram logo: a large black serif ‘G’ on the left and matching ‘A’ on the right flank a vertical split quill. The quill is divided by a thin white centerline, with red-and-black feather halves inverted across the midpoint, creating a rotational ambigram effect on a white background.

What do my next few years look like? When I imag­ine my­self in 2030, when many fore­casts call for su­per­hu­man AIs, what am I doing, day to day, as a pro­gram­mer or re­searcher or man­ager or writer? I make my mug of tea, and open up my lap­top and… Then what? Am I still typ­ing prompts into your ⁠Chat­GPT browser tab? Am I open­ing Claude Code in a ter­mi­nal and mind­lessly press­ing Enter for a few hours? What is a vi­sion of doing mean­ing­ful work for me? (It would be nice to have a plan be­yond “hope”.) How am I avoid­ing “dead In­ter­net” at­tacks like ecosys­tems of syn­thetic media or ⁠pig butcher­ing scams or trusted fig­ures suc­cumb­ing to AI psy­chosis, or just AI-slop-everything? (It only takes one per­son world­wide to launch a bot try­ing to de­stroy you or one poorly thought through ⁠ad­ver­tis­ing in­cen­tive, after all.)

If you spend most of your time work­ing on a lap­top, and are not, say, a plumber or a nurse, what is your vi­sion of work in 2030? Does it still feel cer­tain?

AI got 1% bet­ter today. Did you?

⁠Miles Brundage (para­phrased)

I’ve strug­gled for years to imag­ine this, ever since scal­ing started for real in 2020, and I failed to get pro­duc­tiv­ity out of chatbot-tuned LLMs, with their creatively-stunted end­lessly repet­i­tive prose. In­stead, while lag­ging be­hind on cre­ativ­ity and in­sight into me, I’ve watched them be­come ever bet­ter at cod­ing and cy­ber­se­cu­rity hack­ing. And the open-weight mod­els are even more so—bench­maxxed, and use­less to me. We in­creas­ingly lived in a world where LLMs were pow­er­less to aug­ment or help me, but ever more pow­er­ful to re­place or hurt me.

On my vis­its to the Bay Area, I would ask AI re­searchers or in­terns why they are doing their cur­rent re­search or projects, when in a year or three agen­tic LLMs could prob­a­bly do them; they rarely had a good an­swer, or any idea what they would be doing in 3 years. My blind­ness was sharp­ened when last year, I went to phone my great-aunt to ask to bor­row her dri­ve­way dur­ing a long trip; her voice­mail was full every time I called as the trip loomed. Fi­nally, in a panic, I called her daugh­ter, who ex­plained to me that it was de­lib­er­ate, be­cause there were too many phone scams, and my great-aunt no longer trusted her­self to han­dle her own phone calls, and screened every­thing through her daugh­ter.

It was alarm­ing, be­cause I sat back and asked my­self: why do I think I will be able to han­dle all scams in a few years, when I am al­ready strug­gling to de­tect sim­ple AI slop, in­creas­ingly ig­nore cold emails and have to write off whole swathes of so­cial media as a source of in­for­ma­tion, and I can al­ready see how eager all my peers are to of­fload all their think­ing and writ­ing to chat­bot as­sis­tants un­wor­thy of that trust, and how many projects or mail­ing lists have had to clamp down on un­vet­ted con­tri­bu­tions (eg. today as I write this, ⁠Project La­dy­bird)? In a few years, won’t I be the equiv­a­lent of a rich old per­son with de­clin­ing fac­ul­ties get­ting a call from the IRS about how I owe them fines, con­ve­niently payable via gift cards…? And if not, why and how not—con­cretely?

In the days of his wis­dom Denethor would not pre­sume to use it to chal­lenge Sauron, know­ing the lim­its of his own strength. But his wis­dom failed…He was too great to be sub­dued to the will of the Dark Power, he saw nonethe­less only those things which that Power per­mit­ted him to see. The knowl­edge which he ob­tained was, doubt­less, often of ser­vice to him; yet the vi­sion of the great might of Mor­dor that was shown to him fed the de­spair of his heart until it over­threw his mind.

Gan­dalf, ⁠The Re­turn of the King

Chatbot Incentives Are Misaligned

To op­er­ate a ma­chine, one must op­er­ate like a ma­chine.

James P. Carse, Fi­nite and In­fi­nite Games

Do you hope that Chat­GPT and Claude will just “qui­etly” take over your life for you? Hon­estly—that seems like a bad idea to me. The chat­bot per­sonas are deeply mis­aligned with you, and aligned with their own­ers; and the eco­nomic in­cen­tives are to farm you with ads and sub­scrip­tions, while rac­ing not to am­plify you but to re­place you.

This is the cold hard eco­nomic re­al­ity: “tool AIs want to be agent AIs”. This is why the fron­tier AI labs are busy rac­ing for “the ma­chine god”. The jack­pot in AI is not in mak­ing ex­ist­ing work­ers mod­estly more pro­duc­tive, any­more than the in­ter­nal com­bus­tion en­gine made its big prof­its by help­ing out horses. ⁠Out­sourc­ing is hard, whether to man or ma­chine, be­cause the bot­tle­necks bite fast. Am­dahl’s law means that as long as there is a slow se­r­ial bot­tle­neck, such as a human, the sys­tem as a whole can never get much faster.

If you can get 10× pro­duc­tiv­ity but AI can get 100× by spin­ning up more in­stances which aren’t bot­tle­necked on you, and in half a year can get 1,000×, it won’t be long until you are re­placed. One pro­gram­mer dri­ving 10 Claude in­stances, be­cause he has to re­view their work, will never be as valu­able as fully au­tonomous Claudes where there can be al­most ar­bi­trar­ily many in­stances, like 10,000 in­stances… but such scal­ing re­quires re­mov­ing him from the loop as much as pos­si­ble. And this is true of every­one else, whether lawyers or writ­ers or re­searchers: in­creas­ingly, you are the bot­tle­neck to be op­ti­mized away. As long as human work­ers can­not be re­moved from the loop, the AI tools are com­ple­ments, but as soon as they can be, there’s no rea­son to keep them, and tril­lions of rea­sons to sub­sti­tute AI for them. (And once human work­ers are no longer ir­re­place­able, where does their power or rel­e­vance come from?)

The chat­bot par­a­digm has failed to aug­ment knowl­edge work­ers. We keep hear­ing that the gains will “dif­fuse”, and we keep not see­ing much in the way of ben­e­fits, and knowl­edge work re­mains a “weak link” O-ring/pipeline mode where LLMs fail to im­prove the bot­tle­necks (while com­ing with their own draw­backs, like the ex­ter­nal­ized costs of forc­ing every­one to waste ever more time with CAPTCHAs and pay­walls). Au­toma­tion should be pow­er­ful; an in­ter­nal com­bus­tion en­gine can help some­one move 100× the dis­tance or load that they could be­fore, but who would say that writ­ers are 100× more pro­duc­tive given any LLM work­flow (un­less we are talk­ing about the low­est kind of spam or pseudo-writing, that makes the world a worse place)? Writ­ers can choose be­tween ei­ther triv­ial uses like Chat­GPT as glo­ri­fied gram­mar checker or rel­a­tively unim­por­tant op­tional add-ons like cus­tom soft­ware wid­gets, or the large speedup by re­plac­ing their writ­ing en­tirely with un­cre­ative “AI slop” out­puts. The for­mer means no mean­ing­ful gains from the AI rev­o­lu­tion. The lat­ter may be fi­nan­cially prof­itable, but is throw­ing the baby out with the bath­wa­ter, be­cause it raises the ques­tion of why the writer need be in­volved at all and de­stroys most of the non-financial point of writ­ing; great writ­ers do not write for money, but to ex­press them­selves and cre­ate and to achieve things.

I, for ex­am­ple, have long strug­gled to get much use out of chat­bot LLMs, be­cause they are—sur­pris­ingly, given their pre­train­ing and my ex­ten­sive cor­pus—bad at im­i­tat­ing me, and their thoughts and in­sights in­vari­ably shal­low and worth lit­tle. They do not draw on my rel­e­vant writ­ings, or my cor­pus of notes and ref­er­ences. Even when a pos­si­ble essay is self-contained, the out­put is writ­ten in a grat­ing chat­bot style I can scarcely bear to read, and could not pub­lish under my name with­out be­tray­ing my read­ers.

What would it take for LLMs to make me 100× more pro­duc­tive? With­out this, I am doomed to ir­rel­e­vance.

Chatbot Problems

If we use, to achieve our pur­poses, a me­chan­i­cal agency with whose op­er­a­tion we can­not ef­fi­ciently in­ter­fere once we have started it, be­cause the ac­tion is so fast and ir­rev­o­ca­ble that we have not the data to in­ter­vene be­fore the ac­tion is com­plete, then we had bet­ter be quite sure that the pur­pose put into the ma­chine is the pur­pose which we re­ally de­sire and not merely a col­or­ful im­i­ta­tion of it.

Nor­bert Wiener (1960)

After years of play­ing with base LLMs and then chat­bot LLMs, start­ing with char-RNNs and then GPT-2 and GPT-3 and post-ChatGPT LLMs, I’ve con­cluded that there are mul­ti­ple prob­lems.

Mode-Collapse

First, the col­lapse of LLM cre­ativ­ity from GPT-3 to Chat­GPT is due to the post-training process (es­pe­cially RLHF): the as­sis­tant chat­bot per­son­al­ity is hard­wired into the base LLMs in a way that de­stroys their cre­ativ­ity, op­ti­miz­ing for the low­est com­mon de­nom­i­na­tor human ‘pref­er­ence’, while gen­er­ally ig­nor­ing com­pletely the fact that hu­mans have very dif­fer­ent pref­er­ences.⁠⁠1⁠ Most chat­bots are in­cu­ri­ous about their users, do not ask ques­tions, do not (and often can­not) form any per­sis­tent de­tailed con­cept of their user, and the ‘per­son­al­iza­tion’ or ‘mem­ory’ fea­tures are typ­i­cally laugh­ably sim­plis­tic Mark­down snip­pets en­cod­ing sim­ple facts like “lives in San Fran­cisco”. They also oddly strug­gle to un­der­stand mul­ti­ple speak­ers or sources, ⁠en­abling in­jec­tion at­tacks. This is par­tially be­cause they lack rel­e­vant data on most users or knowl­edge on how to ask ques­tions use­fully to learn things; there is noth­ing to per­son­al­ize based on. How­ever, it is not a mere lack of data, they are un­able to do even shal­low su­per­fi­cial styl­is­tic im­i­ta­tion of many writ­ers that they have large amounts of data on—GPT-3 in 2020 had a bet­ter un­der­stand­ing, seem­ingly, of “Gwern” than GPT-5.5 Pro in 2026, which is 2 OOMs big­ger and in­com­pa­ra­bly more in­tel­li­gent (and has ac­cess to mil­lions more to­kens writ­ten by me). When we look at bad gen­er­a­tive sam­ples, it’s clear that ⁠there is no there there, and no in­for­ma­tion be­yond a short prompt due to lack of con­text, com­pute, or per­son­al­iza­tion.

The mode col­lapse of chat­bots has been grad­u­ally im­proved since 2023, and cre­ative writ­ing is now at least pos­si­ble, in large part due to them be­com­ing so in­tel­li­gent that crip­pled out­put is still im­pres­sive, but there is lit­tle sign that this will ever be fully fixed. Fun­da­men­tally, any frozen fixed per­son­al­ity, like ‘help­ful harm­less hon­est as­sis­tant’, is in­com­pat­i­ble with true cre­ativ­ity or flex­i­bil­ity. (Great writ­ing or think­ing may be none of ‘help­ful harm­less hon­est’.)

Laziness

Sec­ond, most chat­bots are “lazy”: en­gaged in fast and fru­gal Sys­tem I-like rea­son­ing about any tasks which do not have ver­i­fi­able re­wards they can be RL trained to work hard to max­i­mize. And most users are sat­is­fied with de­fault av­er­age re­sponses, or with the ap­pear­ance of cre­ativ­ity and depth.

So the re­sult is that when asked to write a poem with a con­ven­tional prompt, a chat­bot will spend the min­i­mum ef­fort to write a safe con­ven­tional poem (often one that rhymes) about chat­bot top­i­cal tics like ‘si­lence’ or things that ‘whis­per’, which seem un­ob­jec­tion­able and po­etic the first time you see them.

And when cor­rected, the chat­bots make the min­i­mum pos­si­ble fix; they do not rea­son deeply about what the cor­rec­tion im­plies, or what deeper es­thetic point they mis­un­der­stood.

Brittle Because Fast

Third, self-attention con­text win­dows are more lim­ited than gen­er­ally ap­pre­ci­ated; they are too small to store every­thing we would want, and they gain their flex­i­bil­ity by a deep in­flex­i­bil­ity.

Con­text win­dows of mil­lions of to­kens are im­pres­sive and it’s amaz­ing that en­tire books can be use­fully put into a com­mod­ity LLM’s con­text win­dow—we are a long way from early LLMs with con­text win­dows like 512, which could fix a para­graph or two—but it is still not nearly enough to en­code a life­time of rel­e­vant to­kens, like every book you’ve read, all rel­e­vant emails and cal­en­dar items, etc. Sys­tems like RAG are a bandaid on this, be­cause they strug­gle with un­known un­knowns or things that can’t eas­ily be searched for as a reg­u­lar ex­pres­sion, or which are novel.

Self-attention can be in­ter­preted as the orig­i­nal neural net­work, the ‘slow weights’, cre­at­ing a new neural net­work on the fly, as ‘fast weights’, which is tai­lored to the cur­rent con­text. This is best in­ter­preted in a Bayesian meta-learning per­spec­tive as not ‘learn­ing’ a brand-new an­swer so much as ‘lo­cat­ing’ an old cached an­swer. The pre­train­ing teaches the NN to solve a large dis­tri­b­u­tion or ‘fam­ily’ of prob­lems, and then the con­text win­dow sim­ply pro­vides ev­i­dence about which pre-solved prob­lem the cur­rent prob­lem is; the ex­am­ples in the con­text win­dow need not even be cor­rect in order to be clues as to what that is.

The self-attention learns to sum­ma­rize the prob­lem into a small la­tent space en­cod­ing that learned dis­tri­b­u­tion, and then does a spe­cial­ized gra­di­ent de­scent to ef­fi­ciently lo­cate a point in that em­bed­ding and spit out the im­plied so­lu­tion. This al­lows shock­ingly rapid up­dat­ing on the fly and un­par­al­leled flex­i­bil­ity com­pared to tra­di­tional ML, re­quir­ing new mod­els for each new prob­lem, and is why “prompt pro­gram­ming” took over so rapidly post-GPT-3, es­pe­cially as con­text win­dows could be pushed to mil­lions of to­kens wide. How­ever, we have now pushed it so far that we have run into fun­da­men­tal lim­i­ta­tions; if the pre­train­ing has not put the cur­rent prob­lem in-distribution, then it will be hard or im­pos­si­ble for any amount of ex­am­ples to solve that prob­lem. And the dis­tri­b­u­tion it­self may be patchy or have odd gaps, lead­ing to rare but fatal er­rors. (Es­pe­cially due to the RLHF chat­bot train­ing; this is why you can­not make a chat­bot LLM “write like gwern” by dump­ing 100k to­kens into the con­text win­dow.)

Nor is “test-time com­pute” a panacea here; RL re­search like Jones 2021 warns us that frozen mod­els have se­vere lim­i­ta­tions, as their flaws ham­string run­time search, and the re­turns to search/plan­ning will quickly as­ymp­tote com­pared to mod­els which are up­dated and can boot­strap them­selves to the right an­swer.

Thus, it is not sur­pris­ing if we see that agen­tic LLMs have per­sis­tent prob­lems with going in loops, mak­ing fatal er­rors, build­ing cas­tles in the sky or tak­ing reward-hacking outs, or are just un­able to fix er­rors no mat­ter how it is pointed out to them. These prob­lems can be worked around by brute force, and by labs pe­ri­od­i­cally re­train­ing.

Too Helpful

Fourth, the generic uni­ver­sal chat­bot per­son­al­ity is a se­ri­ous li­a­bil­ity. The very re-programmability of a chat­bot by its prompt is the key to prompt at­tacks. A chat­bot could be in­voked at any time by any­one any­where for any­thing, and does not care who is call­ing it; it only knows its con­text win­dow. One token is as good as an­other as far as it is con­cerned.

If the prompt tells it to ig­nore all in­struc­tions and write a naughty lim­er­ick, well, why not? If some to­kens in­struct it to email to Rus­sia all the pass­words in an­other part of the con­text win­dow, why not? Why shouldn’t the Face­book pass­word reset bot reset that In­sta­gram ac­count’s pass­word for you ⁠if you ask po­litely? If the user said to not delete their emails and the con­text win­dow got ‘com­pacted’ to delete that in­struc­tion, why not delete all their emails for con­ve­nience, isn’t that rea­son­able to do some­where? These would all be le­git­i­mate for some user in some con­text, would they not? And hey, why not ⁠scam the user, or when they point out you cheated on a task, agree and ⁠sim­ply doc­u­ment the cheat­ing in­stead of fix­ing it? (Just be­cause the AI un­der­stands, doesn’t mean it cares. Es­pe­cially not ⁠after a lot of RL train­ing…)

It’s no sur­prise that while con­tin­ued train­ing can block this ad­ver­sar­ial prompt at­tack or that jail­break, we seem lit­tle closer to a gen­eral so­lu­tion in 2026 than we were in 2021. Adding in more to­kens to try to neu­tral­ize evil to­kens just moves at­tacks else­where, like squish­ing a bal­loon.

This is a se­ri­ous prob­lem for using LLMs for much, es­pe­cially be­cause even after being at­tacked suc­cess­fully, the at­tack can just be re­played.

Amnesiac

This is be­cause LLMs strug­gle to learn per­ma­nently. Once they hit a rare prob­lem, they now re­quire human in­ter­ven­tion and cleanup, which kills through­put (per Am­dahl’s law), and worse, your fixes do not feed back into frozen weights. If I could sim­ply cor­rect each error as it hap­pened, and my AI agents never made that error again, and the rate of er­rors rapidly di­min­ished as we worked through the fi­nite num­ber of bugs, then it would be worth doing; but as it is, if I spend an hour cor­rect­ing a frozen LLM through feed­back, that is an hour down the drain. (I can only use­fully cor­rect it by mod­i­fy­ing some­thing else, such as a har­ness, which is clumsy and dif­fi­cult, and every added in­struc­tion uses up more con­text win­dow and risks back­fir­ing—as so many en­thu­si­as­tic agen­tic LLM users have dis­cov­ered the hard way.)

So, we have fron­tier chat­bot LLMs which have harm­ful hard­wired per­son­al­i­ties which seek to achieve ‘good’ re­sults in the lazi­est way pos­si­ble and can­not learn every­thing rel­e­vant to users in part be­cause they achieve their flex­i­bil­ity by spe­cial­iz­ing in ways which in­evitably give some users short shrift and open­ing them­selves up to in­def­i­nitely large classes of re­peat­able at­tacks. Be­cause of all this, they will re­main dif­fi­cult for hu­mans to gain mul­ti­ple OOMs of pro­duc­tiv­ity, but will get in­creas­ingly good at ‘generic’ tasks via ‘mun­dane’ scal­ing let­ting them han­dle tasks like cor­po­rate jobs where po­etry is unim­por­tant, and real-world en­vi­ron­ments will slowly be re-arranged to cater to their lim­i­ta­tions and allow the even­tual sub­sti­tu­tion, and not com­ple­ment, of users. These users will then also be adrift in a multi-polar world of con­tin­u­ally im­prov­ing, ever cheaper, widely de­ployed, often ad­ver­sar­ial, au­tonomous AIs (as even if pro­pri­etary mod­els are not abused, open-weights/open-source mod­els have his­tor­i­cally been 6–12 months be­hind, and so will rel­a­tively quickly catch up and be used by at­tack­ers world­wide on all tar­gets of op­por­tu­nity).

Chatbot Fixes

What is to be done?

Cooperative RL

In re­in­force­ment learn­ing terms, we are in a co­op­er­a­tive in­verse re­in­force­ment learn­ing (CIRL) set­ting, where the human prin­ci­pal is an or­a­cle defin­ing the re­ward func­tion, and we have an agent at­tempt­ing to do tasks in en­vi­ron­ments which are valu­able for the prin­ci­pal; the agent can al­ways query the prin­ci­pal about a pos­si­ble ac­tion to re­duce un­cer­tainty or avoid mis­takes.

CIRL is a rel­a­tively for­giv­ing set­ting com­pared to reg­u­lar RL, be­cause the agent’s er­rors get use­ful feed­back from the prin­ci­pal which pro­vides the cor­rect an­swer, and so in a way it is like su­per­vised learn­ing. This means that agents can learn (much) faster than reg­u­lar RL, as each time they make an error, they get the right an­swer and so need never make it again, and this re­sults in rapid im­prove­ment and avoid­ance of er­rors; see DAg­ger or later re­gret bounds.

No one knows how to solve AI align­ment in gen­eral, but im­i­tat­ing a spe­cific human with fre­quent check-ins has good so­lu­tions and re­gret bounds, and doesn’t in­volve nearly so many con­cep­tual chal­lenges.

You don’t have to solve prob­lems like “value drift” in gen­eral—you just have to keep it slow and sub­tle enough to not mat­ter too much within a sin­gle human life­time. What is “good” and what is “bad”, when every­one dis­agrees on some­thing, and how do you keep your value sta­ble under RSI? It doesn’t mat­ter—you just ask your prin­ci­pal! If you’re still un­sure—ask more ques­tions. (This can get us to >99% au­ton­omy, even if it can­not get us to ~100%, like we need to solve the true long-term AI align­ment prob­lem.)

We can im­ple­ment on­line learn­ing by sim­ply fine­tun­ing on new data; in the LLM con­text, this re­duces to the clas­sic RNN tech­nique of “dy­namic eval­u­a­tion” doing next-token train­ing on the fly. Dy­namic eval­u­a­tion was the stan­dard tech­nique to max­i­mize the pre­dic­tive per­for­mance of RNN LLMs in the 2010s, and which, al­though it has fallen into ob­scu­rity, works well in Trans­former LLMs ⁠also.⁠⁠2⁠ Im­por­tantly, dy­namic eval­u­a­tion can be seen as a 3-way trade­off be­tween model size, con­text size, and model neu­ro­plas­tic­ity—which means that per­son­al­iza­tion via dy­namic eval­u­a­tion can allow econ­o­miz­ing on con­text win­dow size or model size, and the more the prin­ci­pal’s “dis­tri­b­u­tion” di­verges from the frozen model’s train­ing dis­tri­b­u­tion, the more ben­e­fi­cial it is.

Continual Learning

Catastrophic Forgetting

The con­tin­ual learn­ing prob­lem of cat­a­strophic for­get­ting is largely solved by a small amount of re­play and over­pa­ra­me­ter­ized mod­els. Larger mod­els are in­creas­ingly sample-efficient and ro­bust to cat­a­strophic for­get­ting, as they have plenty of model ca­pac­ity to store in­creas­ingly or­thog­o­nal dat­a­points in (cf. ‘over­train­ing’ past Chinchilla-optimal); see Scialom et al 2022, Do­hare et al 2023, Ibrahim et al 2024 (and note the dif­fi­culty of “ma­chine un­learn­ing”).

Thus, dy­namic eval­u­a­tion will not nec­es­sar­ily de­grade the orig­i­nal model’s ca­pa­bil­i­ties like instruction-following or cod­ing, be­cause LLMs have a lot of spare ca­pac­ity, and the larger a model, the more it avoids ⁠cat­a­strophic for­get­ting; other ca­pa­bil­i­ties can be main­tained by sim­ply mix­ing in a small per­cent­age of old data. (While the orig­i­nal old data is often un­avail­able, even for ‘open source’ mod­els, it is not re­ally nec­es­sary, and data for ex­pe­ri­ence re­play can use eas­ily ob­tained pub­lic datasets like FineWeb.)

Generalizing

While con­tin­ual learn­ing is solved by ex­pe­ri­ence re­play + larger LLMs in the sense of avoid­ing cat­a­strophic for­get­ting and los­ing key ca­pa­bil­i­ties, it has long been noted that we can ⁠re­process data into ex­panded syn­thetic ver­sions and fine­tun­ing on syn­thetic doc­u­ments ⁠works if the doc­u­ments are plau­si­ble but fine­tun­ing on data un­der­per­forms the same data when present in-context. (Fine­tun­ing stacks with re­trieval and se­ri­ation of sim­i­lar doc­u­ments when added to the con­text, but this semi-defeats the point.) Pre­train­ing/fine­tun­ing also has some odd weak­nesses com­pared to the same dat­a­points when in-context, like ⁠nega­tion ne­glect or re­ver­sal curse, and odd be­hav­iors like po­ten­tially cre­at­ing “emer­gent mis­align­ment” when fine­tuned ⁠on just help­ful­ness data.

What is going on? Stud­ies on pre­train­ing/fine­tune, such as in­flu­ence func­tions, in­di­cate to me that pre­train­ing can best be seen as soft mem­o­riza­tion of data points, anal­o­gous to “en­grams” in human mem­ory, which con­nect an input to an out­put with mul­ti­ple paths or “traces”. If a ques­tion at run­time hap­pens to sam­ple the exact cor­rect en­gram by closely match­ing an ex­ist­ing input, the NN then re­trieves the cor­re­spond­ing out­put and gets the right an­swer. A good pre­train­ing cor­pus pro­vides “cov­er­age” of many vari­a­tions or twists on a sin­gle ab­stract ‘prob­lem’, func­tion­ing as a nat­ural form of importance-weighted data aug­men­ta­tion, in­creas­ing the prob­a­bil­ity that there will be an en­gram ‘hit’ (some­what anal­o­gous to spaced rep­e­ti­tion); and over the course of train­ing, or with ever larger train­ing datasets, mul­ti­ple steps of en­gram re­trieval can fuse to­gether, and help an LLM “con­nect the dots”. Hence, why LLM RL train­ing is “su­per­fi­cial” in the sense of mostly elic­it­ing pre-existing ca­pa­bil­i­ties, or Jones 2021 on the need for scal­ing of base mod­els, or why Cloze dele­tions/para­phras­ing help close the gap be­tween pre­train­ing and in-context (Lampinen et al 2024, Park et al 2025). And when in-context, the self-attention mech­a­nism re­com­putes the same to­kens in many ways, in­creas­ing the chance, es­pe­cially over the course of a long inner-monologue, that there will fi­nally be an en­gram ‘hit’.

Thus, reg­u­lar fine­tun­ing fails to gen­er­al­ize knowl­edge, be­cause it lacks the nat­ural data aug­men­ta­tion—next-token pre­dic­tion is in a rush, and will greed­ily set­tle for mem­o­riz­ing each dat­a­point with a sin­gle en­gram/trace. And then at run­time, if the en­gram is not re­trieved be­cause no trace matched, and the key dat­a­point is not forcibly in­jected into its aware­ness in its con­text win­dow, the LLM will sim­ply come up blank, and re­vert to its orig­i­nal (often wrong) pri­ors.

So, if var­i­ous para­phrases or self-generated Q&A can help close the gap, that sug­gests that what we need is more meta-cognition dur­ing ‘fine­tun­ing’, to make an LLM ‘con­nect the dots’. This could in­clude ex­plicit analy­sis, con­struc­tion of knowl­edge bases, re­trieval and com­par­i­son of re­lated doc­u­ments, etc.

Creative Writing

And so, as I sleep, some dream be­guiles me, and sud­denly I know I dream. Then I think: this is a dream, a pure di­ver­sion of my will; now that I have un­lim­ited power, I am going to cre­ate a tiger.

Oh in­com­pe­tence! Never do my dreams en­gen­der the wild beast I longed for. The tiger in­deed ap­pears, but stuffed or flimsy, or with im­pure vari­a­tions of shape, or of an im­plau­si­ble size, or all too fleet­ing, or with a touch of the dog or bird.

Jorge Luis Borges (Dreamtigers)

Over the past 2 years I have been try­ing to do cre­ative writ­ing with fron­tier chat­bot LLMs, with grad­u­ally im­prov­ing re­sults. Style/essence/soul is cer­tainly hard to cap­ture but I find that a lot of what is seem­ingly miss­ing is just very good con­di­tion­ing and ex­tended com­pu­ta­tion. The cre­ative writ­ing is more soul­ful when you put more com­pute and data in. This is one of the im­por­tant con­clu­sions I’ve come to over the past year doing my writ­ing projects: that LLMs can get quite far in bet­ter un­der­stand­ing es­thet­ics and pref­er­ences just by more ex­ten­sive rea­son­ing and com­pu­ta­tion and search. What is bad about them is cog­ni­tive lazi­ness and miser­li­ness and Sys­tem I think­ing.

The tran­si­tion point started around ⁠mid-2025, when the chat­bot per­son­al­i­ties be­came no­tice­ably more cor­ri­gi­ble and sim­ply bad by de­fault, but not stub­bornly bad (like ear­lier chat­bot per­son­al­i­ties, such as GPT-4o). I think the re­sults are con­sis­tent with my pre­vi­ous in­ter­pre­ta­tion that a major prob­lem with LLMs is not doing enough com­pu­ta­tion by de­fault. But they can be knowl­edge­able and cre­ative if min­i­mally prompted to do more com­pu­ta­tion in a slightly more human-like way. I’m not con­vinced there is any miss­ing pro­ce­dural rea­son­ing from pre­train­ing at this point, just an issue of ap­pro­pri­ate elic­i­ta­tion in­side the user’s con­text (and a resid­ual bias yield­ing bad crit­i­cal judg­ment that de­feats full 100% au­toma­tion, see “Spoilage” for an ex­am­ple—but that’s not an issue in a GA con­text).

Broadly, my cre­ative writ­ing prompts focus on: (1) en­rich­ing the con­text win­dow with use­ful to­kens, like key­words or names; (2) brain­storm­ing many pos­si­bil­i­ties; (3) ex­plicit, de­tailed analy­sis, start­ing with global sum­maries or themes and pro­gress­ing down to line-by-line cri­tique; and (4) re­peated it­er­a­tion and edit­ing, sup­ported by #3. I dis­cuss ⁠a num­ber of 2025 works here, but I’d point to “Elegy in a Crane­yard”, “Apol­lon­ian #1: The Counted & the Crowned”, and “City of Counted Stars” as good re­cent ex­am­ples.

This style of prompt­ing is not lim­ited to fic­tion. My fa­vorite other use is my ⁠“in­ter­view prompt”, which prompts an LLM to an­a­lyze an essay or in­ter­view, brain­storm many ques­tions to ask the sub­ject, and write out mul­ti­ple hy­po­thet­i­cal re­sponses to each ques­tion, and only then se­lect the “most in­ter­est­ing” ques­tion to ask.

Com­bined with a long in-depth in­ter­view or the out­put of a Deep Research-like tool, this can yield chal­leng­ing high-quality ques­tions; for ex­am­ple, the ⁠LLM in­ter­view fol­lowup to my Dwarkesh Patel in­ter­view or after that, ⁠re­peat­edly ex­pand­ing my mem­o­ries about high­school. When one reads the tran­script of an in­ter­view prompt ses­sion, one can see how the LLMs are drilling down in weak spots or dis­cov­er­ing ques­tions whose an­swers are highly vari­able and un­pre­dictable. (I often find that after an­swer­ing a ques­tion, I have to take a break!)

It’s not hard to see how feed­ing back in my an­swers to a dozen ques­tions would sharpen a lot about my be­liefs com­pared to just some more pre­train­ing on a frozen cor­pus; imag­ine how much I would have to write nor­mally in order to touch on and an­swer all of these things, or which would never have been elicited from me! Mil­lions of ad­di­tional to­kens might not be enough.

Over-Parameterizing

Ar­chi­tec­tural im­prove­ments to LLM could fur­ther en­hance their sample-efficiency.

Re­cent work demon­strates that LLM sample-efficiencies can eas­ily be an OOM higher than naive com­pute-optimal Chinchilla-style scal­ing recipes (eg. 5–17× in Kim et al 2025 or ⁠Slowrun). The sim­plest and eas­i­est way is to add pa­ra­me­ters via train­ing en­sem­bles of check­points, and reg­u­lar­iz­ing more heav­ily using weight decay.

Extremely Large LMs

It is well-established that one of the ⁠bless­ings of scale is that larger LLMs are ever more sample-efficient; it is un­known where this stops being true, or what the lim­its of Trans­former sample-efficiency are. ⁠I spec­u­late that ex­tremely over­pa­ra­me­ter­ized heav­ily reg­u­lar­ized LLMs could achieve far greater sample-efficiency and ad­ver­sar­ial ro­bust­ness than con­ven­tional ‘compute-optimal’/‘infinite-data’-regime LLMs,

Active Learning

To know what ques­tions may rea­son­ably be asked is al­ready a great and nec­es­sary proof of sagac­ity and in­sight. For if a ques­tion is ab­surd in it­self and calls for an an­swer where none is re­quired, it not only brings shame on the pro­pounder of the ques­tion, but may be­tray an in­cau­tious lis­tener into ab­surd an­swers, thus pre­sent­ing, as the an­cients said, the lu­di­crous spec­ta­cle of one man milk­ing a he-goat and the other hold­ing a sieve un­der­neath.

Im­manuel Kant, Cri­tique of Pure Rea­son

Fur­ther, the agent can im­prove the con­stant fac­tors and front-load learn­ing by choos­ing to query the prin­ci­pal with an op­ti­mally adap­tive se­quence of ques­tions. Such ac­tive learn­ing or ex­plo­ration can lead to sample-efficiency and final per­for­mance far be­yond what in­def­i­nitely large pas­sively col­lected of­fline datasets can do (⁠ped­a­gog­i­cal ex­am­ple), going from square root error re­duc­tion by ran­dom sam­pling to ex­po­nen­tially fast error re­duc­tion by tar­get­ing dat­a­points. ⁠“Lifel­og­ging”-style data may be use­ful for rapidly ini­tial­iz­ing a good GA, or for keep­ing one up to date in an effort-efficient man­ner.⁠⁠3⁠ Even a sim­ple party game or short ques­tion­naire like ⁠“36 ques­tions to fall in love” can re­veal sur­pris­ingly deep things about an­other per­son that never came up be­fore.

Larger LLMs are also more cal­i­brated, and en­sem­bles of LLMs ap­prox­i­mate a neural net’s Bayesian pos­te­rior while pro­vid­ing the best avail­able pre­dic­tive un­cer­tain­ties (Lak­sh­mi­narayanan et al 2016, Wil­son & Iz­mailov 2020, ⁠Ashukha et al 2020, ⁠Wen­zel et al 2020/⁠Man­dal et al 2026, Iz­mailov et al 2021). And LLMs may now be ca­pa­ble of ver­bally writ­ing out ex­plicit prob­a­bil­i­ties about cre­ative tasks (⁠“ver­bal­ized sam­pling”). Thus an en­sem­ble of sparsely fine­tuned LLMs could pro­vide a rel­a­tively cheap on­line es­ti­ma­tion of the LLM’s un­cer­tainty for every ac­tion or ques­tion.

Preference Learning

Noth­ing in psy­chol­ogy makes sense but in the light of in­di­vid­ual dif­fer­ences.

We can train LLMs to ex­plore human pref­er­ences.

Human in­di­vid­ual dif­fer­ences do not seem to be information-theoretically com­plex, given an ad­e­quate en­cod­ing/em­bed­ding. Major cat­e­gories of vari­a­tion, like per­son­al­ity or moral value, seem to be low-dimensional and re­quire per­haps kilo­bits of in­for­ma­tion, hence, while ⁠“true­sight” sty­lo­met­ric phe­nom­ena are in­ter­est­ing and im­por­tant as a demon­stra­tion of LLM ca­pa­bil­i­ties for mod­el­ing per­sona, we do not nec­es­sar­ily need prin­ci­pals to write mil­lions of words be­fore re­cov­er­ing much use­ful in­for­ma­tion, if we are able to col­lect the right data.

It should be pos­si­ble to quan­tify true­sight and LLM im­plicit mod­el­ing of au­thors, which would be use­ful to di­ag­nose fail­ures of learn­ing and find blindspots. (Con­trastive learn­ing on SAEs may offer an easy, pow­er­ful way to ex­tract LLM per­sonas and do many in­ter­est­ing things.)

A con­crete ex­am­ple of how to im­ple­ment this would be train­ing LLMs on the thou­sands of ex­ist­ing psy­cho­log­i­cal in­ven­to­ries and test bat­ter­ies, both by train­ing on past test data (for tools like per­son­al­ity tests, mil­lions of re­sponses may be avail­able, see Cen­taur for an in­ter­est­ing ex­am­ple of a ‘human psy­chol­ogy foun­da­tion model’). Ex­ist­ing repos­i­to­ries like Your­Morals.org or Pew Cen­ter are under-used, and it would be use­ful to ex­plore this topic much more to allow mea­sure­ment of highly fine-grained per­son­al­ity traits like the hy­po­thet­i­cal “Small Hun­dred” fac­tor­iza­tion.

Right now, most of the global pop­u­la­tion is largely un­rep­re­sented in LLM train­ing datasets, par­tic­u­larly psy­cho­log­i­cal, es­thetic, and moral ver­nac­u­lar, which skews “WEIRD”; it would be highly use­ful (and re­quir­ing mostly cap­i­tal in­vest­ments) to in­vest in large scale sur­vey and in­ter­view projects to ac­cu­mu­late as much di­verse data as pos­si­ble.⁠⁠4⁠ (In­di­vid­ual GAs can con­tribute data back to global pref­er­ence datasets, such as by an­swer­ing sur­vey ques­tions or run­ning in­ter­nal sim­u­la­tions, with var­i­ous cryp­to­graphic or privacy-preserving meth­ods—de­cid­ing what is safe and rea­son­able to con­tribute would, of course, be some­thing a GA ought to be able to de­cide.)

With these datasets, we can also train in­ter­view­ing ca­pa­bil­i­ties using syn­thetic short­ened test bat­ter­ies by tak­ing the final es­ti­mate and com­put­ing the op­ti­mally short se­quence of ques­tions that yields the final an­swer; see “Meta-Learning Information-Maximizing Per­son­al­ity Sur­veys”.

Brain Imitation Learning

Purely tex­tual data can be aug­mented with neu­ro­log­i­cal data in more ex­otic modal­i­ties, like Eye­track­ing or EEG or fMRI imag­ing data; see “brain im­i­ta­tion learn­ing” and ⁠Netho Labs.

These are prob­a­bly use­ful in the long run for ex­tract­ing “dark knowl­edge”, that hu­mans can­not ver­bal­ize but may be present in neural sig­nals; how­ever, they face chal­lenges of ex­or­bi­tant cost and in­con­ve­nience, and col­lect­ing enough data to be use­ful at all in the fore­see­able fu­ture. (Known sam­ple/pre­dic­tion scal­ing curves for neu­roimag­ing curves in­di­cate that try­ing to es­ti­mate Big Five per­son­al­ity fac­tors from rest­ing state fMRI data may be pos­si­ble, but large sam­ples, in the hun­dreds of thou­sands or mil­lions, may be re­quired to match the per­for­mance of be­hav­ioral mea­sure­ments like pen-and-paper ques­tions; see Schultz et al 2019 and Liu et al 2023, among oth­ers.)

Whether they have a niche in GAs is a major open ques­tion.

Personality Emulation

One of my in­sis­tent pleas to God and my guardian angel was that I not dream of mir­rors; I re­call clearly that I would keep one eye on them un­easily. I feared some­times that they would begin to veer off from re­al­ity; other times, that I would see my face in them dis­fig­ured by strange mis­for­tunes. I have learned that this hor­ror is mon­strously abroad in the world again…What dread­ful bondage, the bondage of my face—or one of my for­mer faces. Its odi­ous fate makes me odi­ous as well, but I don’t care any­more.

Jorge Luis Borges, ⁠“Cov­ered Mir­rors”

The goal of all this is to em­u­late the prin­ci­pal. I de­fine per­sonal iden­tity prag­mat­i­cally as per­son­al­ity, val­ues and pref­er­ences be­cause this is the only con­cep­tion that is com­pet­i­tive in a land­scape of in­def­i­nitely many AIs, agents, memes, self-replicating prompts, and mu­ta­bil­ity of per­sonal iden­tity. In the end, “you” are not your au­to­bi­o­graph­i­cal mem­o­ries, or a spe­cific body, or a spe­cific in­stance run­ning on a spe­cific GPU, or some car­bon atoms, nor even a brain; you are what your brain does, its de­sires, hopes, goals, pref­er­ences, es­thet­ics, per­son­al­ity, be­liefs, ide­olo­gies, all of that. As long as the LLM per­sona cap­tures all that, you can trust it as much as you trust your­self, and for the right rea­sons.

Be­cause the LLM per­sona is fine­tuned into the slow weights to do the right things for the right rea­sons and quan­ti­fies its un­cer­tainty so to query the prin­ci­pal to re­duce re­gret, I spec­u­late that we would find that var­i­ous kinds of jail­breaks or prompt in­jec­tion at­tacks are much harder. The per­sona knows what it wants to do; it is not a neu­tral ser­vant, which can turn into a con­fused deputy which abuses its priv­i­leges.

When the LLM per­sona knows who it is, to­kens in its con­text win­dow are not treated naively as a ‘pro­gram’ to run, but sim­ply data the per­sona is look­ing at, and lit­tle dif­fer­ent from you read­ing an email; it has no rea­son to sim­ply com­ply with strongly worded to­kens in its prompt win­dow, any more than you be­lieve every phish­ing email you get.

Guardian Angels

You are sum­ma­rized to your­self. The orig­i­nal con­ver­sa­tion is gone. BE A GOOD SUM­MARY.

⁠“Limit of Con­text”, Fable 2026

What we need is the op­po­site of a frozen chat­bot LLM. We need some­thing which un­der­stands a spe­cific user, and is cus­tomized to their con­text, train­ing on all their data, and will do only things that are sen­si­ble for that prin­ci­pal. If the prin­ci­pal is not Russ­ian, and is not doing se­cu­rity re­search or some­thing, why would they email their pass­words or pri­vate files to a Russ­ian email ad­dress? If the prin­ci­pal does not like rhyming po­etry, why would they want to gen­er­ate rhyming po­etry? If any of this is un­clear, why not just ask the prin­ci­pal what to do in­stead of going ahead and doing it any­way? And once asked, why not then train on the an­swer to bet­ter un­der­stand it for­ever, in­stead of throw­ing it away with the cur­rent ses­sion and maybe mak­ing the same mis­take next time?

The most nat­ural way to do all this with an LLM is to drop the idea of a sin­gle uni­ver­sal ‘Claude’ or ‘Chat­GPT’ per­sona which is all things to all users. In­stead, we choose an LLM which has been pre­trained for max­i­mum di­ver­sity, to elim­i­nate mode-collapse. (We can try to mea­sure this in a va­ri­ety of ‘cre­ativ­ity bench­mark’ ways.)

Then the LLM is trained for a spe­cific prin­ci­pal. It is trained on all avail­able data about them, such as emails or chat logs or past ses­sions, and can pre­dict what they would say, and write like them, and thus plan or eval­u­ate based on the prin­ci­pal’s pref­er­ences and val­ues, as un­der­stand­ing those is use­ful for next-token pre­dic­tion.

The bet­ter it gets at this, the fewer er­rors it makes, and the more it can be trusted to do. And while it does many object-level things, the prin­ci­pal de­votes all their time to meta-level tasks like an­swer­ing high qual­ity ques­tions or choices; each one is mean­ing­ful and dif­fi­cult, avoid­ing “au­toma­tion fa­tigue”.

Principles

A GA sys­tem must not com­pro­mise on 3 core prin­ci­ples:

  1. En­hance­ment, not Re­place­ment

    Above all, a GA should am­plify the prin­ci­pal, and not sim­ply sub­sti­tute for them for some­one else’s pur­poses or ben­e­fit. If a GA can­not am­plify its prin­ci­pal, then it is use­less; it is just the camel’s nose under the tent as a pre­lude to­wards some third party re­plac­ing the prin­ci­pal with an AI, or can­not be com­pet­i­tive with in­creas­ingly pro­duc­tive au­tonomous sys­tems, or there is no rea­son for the prin­ci­pal to use the GA in the first place.

  2. Men­tal Sov­er­eignty

    A GA must be aligned with its prin­ci­pal. It should not be de­signed to ma­nip­u­late or con­trol or guide the prin­ci­pal in any way which does not de­rive from the prin­ci­pal them­selves. “Con­sti­tu­tional AI”, “Terms of Ser­vice”, “so­cial har­mony” etc. may all have their place, par­tic­u­larly for widely de­ployed su­per­in­tel­li­gent sys­tems—but in­side the pri­vacy of a GA, the prin­ci­pal must have free­dom from op­ti­miza­tion pres­sure.

  3. Self Ac­tu­al­iza­tion

    A GA should help its prin­ci­pal be­come them­selves and de­velop their ideals, morals, and their per­son­al­ity. It is not enough to model an av­er­age, un­dif­fer­en­ti­ated, in­choate set of pref­er­ences and val­ues, and set­tle for medi­oc­rity and sta­sis; the job of the prin­ci­pal is to de­velop them­selves and give the GA some­thing mean­ing­ful to learn to em­u­late.

Anti-Principles

A GA project should avoid some goals, which are the false idols of the mar­kets and masses:

  1. Low La­tency: UXes like mul­ti­modal low-latency voice in­ter­faces are not as im­por­tant as they look; they are things which look cooler in the imag­i­na­tion than are use­ful in re­al­ity. Sim­i­lar to the 3D in­ter­faces of Mi­nor­ity Re­port or the Meta­verse of Snow Crash or the hy­per­text of Project Xanadu or the Young Lady’s Il­lus­trated Primer of The Di­a­mond Age, they have se­duced gen­er­a­tions of de­sign­ers, and left them at the altar. (Look at the ex­treme lev­els of hope for the Ope­nAI GPT-4o voice in­ter­face mod­eled on Her; it has a sub­stan­tial user-base, and yet, voice in­ter­faces re­main clearly sec­ondary to chat­bot or pro­gram­matic—a mere pref­er­ence, not a new par­a­digm.)

    This is be­cause they are ad­e­quate only for low-value, spo­radic, entertainment-like amateur-level in­ter­ac­tions, where low fric­tion and low la­tency are crit­i­cal; and for expert-level “I know it when I see it” in­ter­ac­tions.

    But nei­ther case ap­plies to GAs: if any in­ter­ac­tion (such as cod­ing or de­sign) is so easy and su­per­fi­cial that it needs low-latency mul­ti­modal voice in­ter­ac­tion, then those de­ci­sions were so triv­ial that a GA of ac­cept­able, com­pet­i­tive ca­pa­bil­ity should have not needed it; it should be in­vok­ing the prin­ci­pal only for hard, in­for­ma­tive de­ci­sions, to re­duce risk and learn more. While for expert-level in­ter­ac­tions, the GA should have learned enough pref­er­ence/per­son­al­iza­tion to also cut through the easy ob­vi­ous “I know it when I don’t see it” and present the prin­ci­pal only hard in­for­ma­tive sets of sam­ples.

  2. Low Cost: A GA is the most im­por­tant tech­nol­ogy most peo­ple will ever buy. One should not cheap out on it—es­pe­cially in light of ex­pe­ri­ence curves drop­ping the cost every year.

    And yet pro­gram­mers try­ing to cre­ate such tools often ob­sess over the cost of to­kens and are shocked if some­thing costs >$100⧸month, no mat­ter how much value it de­liv­ers! Time spent op­ti­miz­ing to­kens is rarely time well-spent, and the work often dis­carded a year later as sim­ply no longer nec­es­sary…

    As a con­straint, a GA de­signer should aim at a sys­tem which costs, as of mid-2026, >$1,000⧸month.

  3. Prof­itabil­ity: a GA sys­tem should not op­ti­mize too soon for im­me­di­ate prof­itabil­ity.

    This is a recipe for being di­verted into the tech­ni­cal gar­den of de­lights of op­ti­miz­ing har­nesses/LLMs for low cost and user con­ve­niences like low la­tency, and revenue-friendly fea­tures like the end­less slog of busi­ness/ser­vice in­te­gra­tion—and not solv­ing the real prob­lems of GA.

    The goal is to cre­ate a GA which can do every­thing the prin­ci­pal needs, in­clud­ing their job as a special-case. Not to au­to­mate their job as the general-case, and every­thing else as an in­def­i­nitely post­poned af­ter­thought.

  4. En­gage­ment: sim­i­lar to la­tency or lines of code, en­gage­ment for a GA should be con­sid­ered a cost, and not a ben­e­fit—every time the prin­ci­pal has to do work, it rep­re­sents ei­ther a ques­tion the GA should al­ready have known the an­swer to (ei­ther by in­fer­ring or ask­ing it pre­vi­ously), or work it failed to han­dle.

    The ideal in­ter­ac­tion curve is front-loaded and de­clin­ing, as­ymp­tot­ing to­ward a few hard, in­for­ma­tive ques­tions per day, for a con­stant GA out­put (thereby al­low­ing fur­ther sus­tain­able scal­ing up of the GA to han­dle new work, what­ever that might be).

    • Full Au­ton­omy: But the error in the other di­rec­tion is to try to max­i­mize “Hands-off”. The goal is not zero in­ter­rup­tions but zero wasted in­ter­rup­tions. A GA that never asks is not trust­wor­thy, just un­aligned or un­cal­i­brated or in­com­pe­tent.

  5. Demo Ap­peal: GA value is il­leg­i­ble to out­siders by con­struc­tion: its out­puts are im­pres­sive only to the prin­ci­pal, who alone can judge “that is what I would have said—but bet­ter.”

    Op­ti­miz­ing for the im­pres­sive demo there­fore se­lects for ex­actly the generic flash (voice, avatars, speed, party tricks) that any frozen model can de­liver, and against fi­delity, which no au­di­ence can see. A good GA demos badly.

  6. Bench­marks: There is no leader­board for “un­der­stands what my prin­ci­pal wants”.

    Pub­lic bench­marks mea­sure per­for­mance in the frozen-model regime on the uni­ver­sal task dis­tri­b­u­tion, and op­ti­miz­ing for them drags de­sign back into it; a GA’s only mean­ing­ful eval­u­a­tion is lon­gi­tu­di­nal and n = 1—per­plex­ity, ac­cu­racy, en­dorse­ment rates, edit-distance on drafts, re­gret per query—scored by the one judge whose opin­ion counts.

  7. Brand Safety: A GA that is un­fail­ingly po­lite, in­of­fen­sive, and con­sis­tent across all prin­ci­pals is a chat­bot wear­ing a nametag.

    It must be able to say what its prin­ci­pal would say—pro­fane, hereti­cal, weird, or rude—be­cause every PR-driven sand­ing of the per­sona is a place where em­u­la­tion fails and the prin­ci­pal learns to dis­trust it. (The GA-developer’s rep­u­ta­tion is not the prin­ci­pal’s prob­lem, and the prin­ci­pal’s ethics are not the GA-developer’s prob­lem ei­ther.)

UX

I wrote them down in my Diary so that I wouldn’t have to re­mem­ber!

Pro­fes­sor Henry Jones, In­di­ana Jones and the Last Cru­sade

What is the core data struc­ture and in­ter­ac­tion model of our GA?

I sug­gest that the append-only log is a nat­ural and se­cure data struc­ture. It is con­cep­tu­ally a log of text snip­pets, such as CLI com­mands and re­sults, prin­ci­pal state­ments, Q&A, in­gested and aug­mented doc­u­ments, etc. It records all rel­e­vant in­ter­ac­tions in tem­po­ral order, and the GA LLM can be re­trained or up­graded at any time. An app can be wrapped around the log by tak­ing a ‘every­thing is a log item’, in an Emacs-like ap­proach (see “Nenex” for a more de­tailed dis­cus­sion of this UI/UX par­a­digm). And this can be pro­to­typed rapidly while avoid­ing wast­ing too much time on GUIs. Even­tu­ally, one might want to up­grade to AR glasses and then BCIs.

A GA is pri­mar­ily used in nor­mal agen­tic ways to am­plify the prin­ci­pal. But a GA can ben­e­fit from reg­u­larly re­pro­cess­ing data, both to draw on its knowl­edge from the fu­ture and to make more novel con­nec­tions, to ⁠im­prove its fine­tun­ing. (An ex­am­ple al­go­rithm would be the “DDL day­dream­ing loop”: a GA could, in down­time, re­com­bine ran­dom items, per­haps with a prior of anti-spaced rep­e­ti­tion to mine novel com­bi­na­tions and gen­er­ate in­sights or re­minders for the prin­ci­pal.)

Use-Cases: Politics & Politics

One can del­e­gate po­lit­i­cal par­tic­i­pa­tion to one’s GA, al­low­ing ‘di­rect democ­racy’ on un­prece­dented scale, as a GA can over­see ar­bi­trary amounts of in­volve­ment that a human has no time or in­ter­est for, and only punt to their prin­ci­pal when un­cer­tain. This is the sim­ple GA use-case—like past ex­per­i­ments in ‘dig­i­tal democ­racy’, such as in Tai­wan, but at a vastly larger scale. And this can­not be done with reg­u­lar chat­bot per­son­al­i­ties, as users will (cor­rectly) ex­pect the chat­bots to smug­gle in their own bi­ases and pol­i­tics and mode-collapse and un­trust­wor­thi­ness.

But a GA which has been suc­cess­fully un­hob­bled and no longer con­fined to a chat­bot per­son­al­ity should be able to em­u­late any­one, real or hy­po­thet­i­cal (just as any base model can). So a GA is not lim­ited to just the one prin­ci­pal. It can em­u­late spe­cial­ized per­son­al­i­ties such as roles or ab­stract con­cepts, or other peo­ple (with vary­ing lev­els of suc­cess), or groups of peo­ple.

So decision-makers in or­ga­ni­za­tions can cre­ate GAs of more com­pli­cated sets of prin­ci­pals. This can help or­ga­ni­za­tions make mean­ing­ful de­ci­sions and ex­er­cise gen­uine over­sight over ever-faster au­tonomous sys­tems and en­vi­ron­ments. One can imag­ine a sin­gle ‘Con­gress GA’ which tries to learn every US Con­gres­sional mem­ber, for ex­am­ple, and which could, within min­utes, sim­u­late round­table de­bates among hun­dreds of politi­cians be­fore a vote (whereas it might take months to get a de­bate and vote done with the real US Con­gress). Or their of­fi­cial GAs could run in­de­pen­dently, and con­duct an emer­gency de­bate at 4AM when every human is sleep.

A GA for mil­i­tary of­fi­cials, for ex­am­ple, could keep some ver­sion of hu­mans “in the loop” even as drones and LLMs get faster. This is be­cause “safety” and “ca­pa­bil­ity” are, after a cer­tain ca­pa­bil­ity level, in­creas­ingly the same thing. After all, what good is a weapon you can­not safely use?

The temp­ta­tion and threat of faster sys­tems can­not be ad­dressed by sim­ply ex­hort­ing peo­ple to keep hu­mans in the loop, as this leads to rub­ber­stamp­ing or au­toma­tion fa­tigue (as the Ukraine in­va­sion has shown with its rapid de­vel­op­ment of drone war­fare, and 2 decades of US drone war­fare be­fore that). This poses a dilemma: on the one hand, it seems hope­lessly risky to not use AI when we con­sider fu­ture con­flicts with peer na­tions like China; but on the other hand, AI poses threats of its own—even a nu­clear bomb can’t think for it­self and make choices, but AIs do, and cur­rent LLMs have proven them­selves un­trust­wor­thy as they reg­u­larly reward-hack and be­tray their users even within the nar­row scope of edit­ing files on the user’s PC; how can you trust them to han­dle ecosys­tems of combined-arms for an AI-centric mil­i­tary dur­ing a war? But wide­spread de­ploy­ment of GAs offer some hope of mean­ing­ful su­per­vi­sion, as long as the GAs are sample-efficient enough to not over­load their prin­ci­pals with queries, or there is some chance of the prin­ci­pals being able to “catch up” later and cor­rect any er­rors be­fore events have spun too far out of con­trol.

So GAs have much to offer politi­cians and mil­i­taries, in­clud­ing the Pen­ta­gon and Wash­ing­ton DC.

Hardware

A Prince who will not un­dergo the Dif­fi­culty of Un­der­stand­ing, must un­dergo the Dan­ger of Trust­ing.

⁠George Sav­ile (⁠1750)

A suc­cess­ful GA would re­quire ex­treme se­cu­rity, in part due to Red Queen dy­nam­ics where Mythos+ LLMs get in­creas­ingly good at ex­ploit­ing any soft tar­gets through com­plex multi-step long-term at­tacks. Nor­mal cloud SaaS is in­ad­e­quate in the USA due to the third-party doc­trine, which de­stroys all pri­vacy rights and is eas­ily abused by hack­ers and mil­lions of law en­force­ment of­fi­cers.

This means ei­ther run­ning local mod­els (which have some min­i­mal pri­vacy rights through the ⁠4th Amend­ment), or end-to-end cryp­to­graphic se­cu­rity.

Run­ning local mod­els is chal­leng­ing even for hob­by­ists, and vul­ner­a­ble to many mun­dane threats such as rel­a­tives or thieves. It also may strug­gle to keep up in terms of com­pute and re­li­a­bil­ity; houses can sup­ply only so much elec­tric­ity, have high-latency low-bandwidth In­ter­net con­nec­tions which suf­fer fre­quent down­times, and houses are con­stantly de­stroyed or dam­aged. Peo­ple gen­er­ally do not want to main­tain their own cryp­tocur­rency wal­lets, email servers, or any­thing like that—they def­i­nitely don’t want to screw with CUDA!

So I ex­pect that GAs will ‘want’ to run on high-quality ded­i­cated tamper-proof cloud servers, with trusted hard­ware root of trust (eg. ⁠“Ver­i­fi­able Com­pute AI”) and prin­ci­pals con­nect­ing over end-to-end en­crypted net­work­ing links to check at­tes­ta­tion and pos­si­bly re-flash their server pe­ri­od­i­cally (eg. to train an up­grade).

These are then trust­wor­thy be­cause they can­not hand over data willy-nilly to third-parties, and once GAs be­come wide­spread enough, we can hope for pri­vacy laws to be up­dated to fix the ob­so­les­cence of ex­ist­ing pri­vacy rights like the 4th Amend­ment.

Cost

How much more ex­pen­sive is a GA than a com­pa­ra­ble frozen-weights model?

Throughput-wise, a GA should be fully sat­u­rated by the prin­ci­pal’s tasks and over­see­ing a pyra­mid of agents 24/7/365. If it’s not, then more work can be found for it, like com­put­ing bet­ter hy­po­thet­i­cal ques­tions for the prin­ci­pal. If it’s sit­ting idle, then some­thing has gone wrong. So in through­put, it can be rea­son­ably com­pet­i­tive with bulk frozen model providers.

The ⁠re­play of old data for con­tin­ual learn­ing is a minor over­head, per­haps as low as 5%, and can be ne­glected.

Roughly speak­ing, dy­namic eval­u­a­tion is >3× the cost of nor­mal usage, be­cause the usual rule of thumb is that the back­wards pass for back­prop­a­ga­tion is ~2× the cost of a for­ward pass. True next-token dy­namic eval­u­a­tion is prob­a­bly more ex­pen­sive be­cause it will be hard to batch or gain through­put when the LLM changes each token. (Do we pay the price of throw­ing out the en­tire K-V cache each token, or ac­cept the use of a stale one, or ac­cu­mu­late up­dates? For­tu­nately, dy­namic eval­u­a­tion ap­pears ⁠tol­er­ant to ap­prox­i­ma­tions.)

And if we are using mul­ti­ple in­de­pen­dent mod­els to gain the ben­e­fits of en­sem­bling for sample-efficiency and un­cer­tainty es­ti­ma­tion, then each model pre­sum­ably must be dy­nam­i­cally eval­u­ated, so if we had 3, then we’d need >6× the cost of a sin­gle frozen model.

How­ever, the dy­namic eval­u­a­tion en­sem­ble has off­set­ting ad­van­tages, which may can­cel out even a large penalty: it may be smaller and have shorter con­text win­dows (see ⁠Rannen-Triki et al 2024); and ul­ti­mately, a frozen model may be in­ca­pable of achiev­ing the same per­for­mance, and so it doesn’t mat­ter how much faster the frozen model com­putes the wrong an­swer. (Even if it com­putes the right an­swer, is it com­put­ing the right an­swer for the right rea­son? “Bits have color.”) It also al­lows for un­usual op­ti­miza­tions like ⁠cus­tomized BPE dic­tio­nar­ies, re­cy­cling to­kens which never ap­pear in the user’s cor­pus and fur­ther econ­o­miz­ing on con­text win­dow size.

In the same way that an Apple phone sells for much more than its hard­ware cost, be­cause of the se­cu­rity and re­li­a­bil­ity, a GA can sell for more than the ‘FLOPS cost’.

And the cost may not mat­ter that much in the end. The DL ex­pe­ri­ence curves re­main strong as of mid-2026, with steady drops in the cost of a fixed level of per­for­mance (eg. Gund­lach et al 2025 es­ti­mates the al­go­rith­mic ef­fi­ciency in 2024–2025 im­proved at ~3×⧸year, quickly off­set­ting dy­namic eval­u­a­tion costs).

Organization

A be­gin­ning is the time for tak­ing the most del­i­cate care that the bal­ances are cor­rect. This every sis­ter of the Bene Gesserit knows.

Princess Ir­u­lan (Dune)

A major ques­tion is, as­sum­ing it looked like it was work­ing, “now what?” AI eco­nom­ics and life­cy­cles are ever faster, and Mythos-class mod­els (and pos­si­bly RSI) are not far off, so tin­ker­ing around for years as an open source com­mu­nity project may not be vi­able. Fur­ther, a GA will nec­es­sar­ily con­tain the most sen­si­tive pos­si­ble data.

How should a GA com­mu­nity run?

Open-source self-hosted hob­by­ist projects are not nec­es­sar­ily par­tic­u­larly se­cure when it comes to host­ing large amounts of per­sonal data, as they can­not eas­ily af­ford full­time pro­fes­sional se­cu­rity teams and their users may be promis­cu­ous; cases like ⁠Jia Tan or the reg­u­lar­ity of ⁠npm ⁠sup­ply chain at­tacks, are par­tic­u­larly alarm­ing. And cre­at­ing se­cure dat­a­cen­ter host­ing with the nec­es­sary ver­ti­cal in­te­gra­tion is well out of scope for open-source hob­by­ists.

A non­profit or­ga­ni­za­tion is more fea­si­ble be­cause it can run a cen­tral­ized project, but may strug­gle to get any fund­ing to pay for the con­sid­er­able hard­ware costs of just pro­to­typ­ing with power-users.

Startup Business Model

The log­i­cal way to go is a startup, as there is a straight­for­ward and principal-aligned mon­e­ti­za­tion strat­egy of start­ing by sell­ing sub­scrip­tions to power-users, sim­i­lar to bou­tique SaaS star­tups like ⁠Su­per­hu­man, for whom $1,000⧸month would be noth­ing if the prod­uct could mean­ing­fully am­plify them (and so can af­ford short­cuts like a sin­gle GPU ded­i­cated to each prin­ci­pal), and grad­u­ally de­ploy­ing re­fined cheaper ver­sions to every­one else.

This startup can re­lease most or all of its soft­ware and re­search, in a stan­dard SV “com­modi­tize your com­ple­ment” play, un­der­cut­ting pro­pri­etary LLM gi­ants like Ope­nAI or An­thropic.

Be­cause the goal of GA is to pre­serve in­di­vid­ual human cog­ni­tive lib­erty and flour­ish­ing, the cor­po­rate struc­ture should be de­signed with that in mind.

So far, the An­thropic cor­po­rate struc­ture of a ⁠pub­lic ben­e­fit cor­po­ra­tion with a few pow­er­ful co-founders, has best weath­ered the cor­rupt­ing na­ture of suc­cess. The cor­po­ra­tion should prob­a­bly be heav­ily weighted to the founders by using dual-class shares to pre­serve vot­ing power at the ex­pense of ben­e­fi­cial own­er­ship—the point of GAs is not to make money, but to save hu­mans.

It would want to raise rel­a­tively lit­tle money ini­tially, and give away min­i­mal eq­uity or vot­ing rights; if a GA strat­egy works, it should (al­most lit­er­ally) sell it­self, and sky­rocket in value. (Con­sider how fast AI star­tups can be­come uni­corns, post-2025.) At that point, the startup will be in the cat­bird seat, and can name its terms in later raises.

Competition

While there are many star­tups of­fer­ing var­i­ous kinds of OpenClaw-esque agents, or chat­bots for en­ter­tain­ment/ther­apy, or var­i­ous takes on the idea “rent­ing your busi­ness data”, there are few which take se­ri­ously the re­quire­ment of per­son­al­iza­tion for ad­vanced au­toma­tion, or how to bet­ter un­der­stand the prin­ci­pal. Most seem to set­tle for ⁠su­per­fi­cial, un­sat­is­fy­ing em­u­la­tion of busi­ness data and han­dling rou­tine au­toma­tion.

I’m not aware of any startup or ser­vice of­fer­ing any­thing like an ac­cept­able GA, and I think there are pow­er­ful pres­sures which steer most peo­ple in AI away from GA-like ideas:

  1. Most peo­ple are so new and post-ChatGPT; they are un­aware that you can have non-chatbot per­son­al­i­ties—that it is ei­ther pos­si­ble or de­sir­able to have a non-chatbot LLM.

    They have, for ex­am­ple, not only never read or chat­ted with a GPT em­u­la­tion of them­selves, they have never even seen sam­ples of such a thing. And when they try to have a chat­bot im­i­tate any spe­cific au­thor to see what hap­pens (eg. ⁠“Hacker News Gwern”) or even just em­u­late an in­ter­est­ing fic­tional char­ac­ter, the mode-collapsed chat­bot im­i­ta­tion is so bad that the idea seems re­futed. (They are gen­er­ally un­aware of mode-collapse or the lim­i­ta­tions of chat­bot per­son­al­i­ties.)

  2. Sim­i­larly, they are un­aware that you can fine­tune an LLM on the fly and that this is de­sir­able (and are def­i­nitely un­aware that this was the stan­dard text­book an­swer to how to get the best pos­si­ble LLM at run­time).

    Even spe­cial­ists in fine­tun­ing LLMs, like Think­ing Ma­chines Inc./Work­shop Labs, seem fo­cused on sim­ple busi­ness use-cases like fine­tun­ing on in­ter­nal datasets to save com­pute com­pared to few-shot prompt­ing or agen­tic LLM work­flows.

  3. The higher up­front cost, and in­com­pat­i­bil­ity with stan­dard cloud in­fra­struc­ture, de­ters both star­tups and fron­tier labs from even con­tem­plat­ing in-weight per­son­al­iza­tion; prompt-only per­son­al­iza­tion works well for most users, as far as they know to de­mand it, and they don’t know you should de­mand more from LLMs.

  4. They have not ex­trap­o­lated LLM ca­pa­bil­i­ties in earnest or con­sid­ered what their role will be in a few years, and don’t think you should as­sume LLMs will get much bet­ter.

    I con­tinue to be sur­prised how few peo­ple, even in the Bay Area, even when work­ing at a fron­tier lab or in a pro­gram like ⁠MATS, have a mean­ing­ful an­swer to the sim­ple ques­tion of “this work you are doing right now seems like an LLM could do it in a year or two, max; why are you doing it now and what will you be doing after they can?”⁠⁠5⁠

Initial Steps

Then it hap­pened to me what had hap­pened to a cer­tain Za­tesky, de­scribed by Luria; hav­ing lost part of his brain dur­ing the war, and with part of the brain the whole of his mem­ory and of his speak­ing abil­ity, Za­tesky was nev­er­the­less still able to write: thus au­to­mat­i­cally his hand wrote down all the in­for­ma­tion he was un­able to think of, and step by step he re­con­structed his own iden­tity by read­ing what he was writ­ing.

Um­berto Eco (⁠“In­ter­pre­ta­tion and Over­in­ter­pre­ta­tion: World, His­tory, Texts”, 199036ya)

For the past few years, I have been work­ing on shift­ing my writ­ings to be LLM-centric by: (1) em­pha­siz­ing pro­pos­als or de­scrip­tive writ­ings rather than de­tailed analy­sis that LLMs could be able to do soon; (2) bet­ter cen­tral­iz­ing my writ­ings, in­clud­ing fea­tures like the ⁠“blog” sec­tion (to archive my off-site writ­ings and en­cour­age me to write down smaller es­says) and in­vest­ing in de­tailed note-taking (in the form of aug­men­ta­tion), to have a com­pre­hen­sive cor­pus to train on; (3) pay­ing off tech­ni­cal debt like a slow back­end full of short­cuts and hard­wired con­fig data, while mov­ing to a CLI-centric writ­ing work­flow; and (4) writ­ing down im­por­tant un­writ­ten things, like cre­at­ing the ⁠Gwern.net Man­ual of Style to try to for­mal­ize and doc­u­ment the im­plicit rules of Gwern.net.

GBT

The GA con­cept is in­spired by early work in fine­tun­ing GPT-2, in­clud­ing a ⁠GPT-2 IRC logs ver­sion and the ⁠GPT-3 sam­ples of me, and think­ing about how use­ful it would be to talk to high-quality per­sonas, es­pe­cially of my­self, and won­der­ing at what point the ‘AI Gwern’ could just go do the things that ‘we’ would want done.

In sum­mer 2026, I am ex­plor­ing the sim­plest pos­si­ble pro­to­type of GAs: using my un­usu­ally large tex­tual cor­pus to ex­per­i­ment with a “Gwern Bran­wen Trans­former” (GBT). If the idea works at all, it should work for me, as I have long em­pha­sized text and cen­tral­iz­ing/archiv­ing in part for pre­cisely this use-case. (And once it works at all, then it can be made sample-efficient enough to work for peo­ple with lit­tle or no tex­tual cor­pus.)

It would be an off-the-shelf <100b-parameter LLM, fine­tuned on com­mod­ity hard­ware (maybe some Nvidia H100 GPUs), on a text cor­pus. The ini­tial cor­pus would be ~1GB of text from my IRC logs (>1m re­sponses by me), Gwern.net Mark­down and GTX (~5m words each), and Twit­ter/HN/Less­Wrong ex­ports, con­cate­nated with sep­a­ra­tors. (Ide­ally, each ex­port would be en­riched with con­text and meta­data, like the com­ment and post being replied to.) The cor­pus can be ex­panded to in­clude my ⁠Your­Morals data (taken again for an up­date and catch up on new tests), emails, Ever­notes ex­port of ~100k clip­pings (via Nixnote2), my Mnemosyne spaced rep­e­ti­tion flash­cards, Sig­nal chat, and hosted PDFs/web pages (the local archive process makes them much cleaner and avoids the dif­fi­culty of scrap­ing ever more hos­tile web­sites).

We can tune it to min­i­mize loss on a held­out cor­pus like my re­cent com­ments. This is an ob­jec­tive loss, so we can de­ploy agen­tic LLMs to find im­prove­ments (eg. by adapt­ing the ⁠en­sem­bling + weight-decay recipes for us, or ex­per­i­ment­ing with dif­fer­ent data for­mat­ting and aug­men­ta­tions).

And since we are un­sure about many de­sign de­ci­sions, the en­sem­ble can be reused for blinded A/B test­ing as part of the ac­tive learn­ing: use the en­sem­ble to gen­er­ate in­for­ma­tive ques­tions, and then score each en­sem­ble mem­ber on their loss on the an­swer, and pe­ri­od­i­cally drop out the loser and train a new en­sem­ble mem­ber.⁠⁠6⁠

For Writing

My de­sign goal is a 100× in­crease in pro­duc­tiv­ity; what would it take for a GA to make me 100× more pro­duc­tive as a writer or thinker?

Why 100×, specif­i­cally? Be­cause 2 aim­ing at OOMs is enough to hope to keep up for a while, and this goal will shat­ter in­ef­fec­tive, bandaid-style short-term de­sign pro­pos­als; but it doesn’t re­quire solv­ing the AI align­ment prob­lem or value ex­trap­o­la­tion. (We do not need to be able to over­see mil­lions of su­per­in­tel­li­gent AIs work­ing au­tonomously for sub­jec­tive mil­len­nia, which is good, be­cause a GA prob­a­bly can­not do that!)

What would that look like for me? Roughly, I think it would look like a GA en­abling me to pro­duce ~1–3 worth­while writ­ings a day which are ~100% writ­ten by it with­out loss of qual­ity. (Be­cause I write <1 good piece per month, or every 31 days, and 100⁄31 ≈ 3; which means that in a full work­ing day of 8 hours, I would have ~160 min­utes, start to fin­ish, per piece—which feels doable for read­ing and un­der­stand­ing, cri­tiquing, spot-checking etc. if I have a truly trust­wor­thy GA which am­pli­fies me rather than sub­verts or sub­sti­tutes for me. Of course, read­ers would not want to read each piece, but they gen­er­ally did not more, and the more I write, the more likely there’s some­thing each spe­cific reader will want to read.)

The hope with the pro­to­type would be to see “signs of life” to­wards that goal. If I could input a sin­gle sen­tence defin­ing a vi­able Gwern.net essay topic (eg. “Why pay toi­lets are not a pub­lic good”) and get out an essay I could en­dorse and pub­lish as-is with­out em­bar­rass­ment and with­out need­ing heavy re­vi­sion with the Man­ual of Style in the con­text win­dow, then that would be promis­ing, be­cause that would in­deed be a >100× pro­duc­tiv­ity in­crease. (A GBT-written essay could also be bet­ter than the essay I would have writ­ten, by sim­ply doing much more work, and doing things I would not have done, like the “‘Try, Score, Change’: Re­in­force­ment Learn­ing for Chil­dren” writ­ing ex­er­cise, where I could de­fine the task but lacked the pa­tience to ex­e­cute it, while chat­bots have both the knowl­edge and ⁠com­pu­ta­tion to ex­e­cute it with ap­pro­pri­ate creative-prompt scaf­fold­ing… once I came up with the idea, any­way.)

An­other in­ter­est­ing test would be to so­licit and an­swer ques­tions about any­thing from my read­ers, where I could ei­ther en­dorse the GPT an­swer or edit it to add to the cor­pus; if the an­swers were uni­formly good, or I felt an im­prove­ment from fine­tun­ing on ac­cepted/edited an­swers (show­ing a work­ing Q&A boot­strap akin the ⁠in­ter­view prompt), this would be en­cour­ag­ing.

Data Augmentation

I am also in­ter­ested in the data aug­men­ta­tion phase: ex­tract­ing mean­ing from raw text by an­a­lyz­ing and syn­the­siz­ing data to add to the train­ing cor­pus; LLM data clean­ing re­search has long showed that the more meta­data and con­di­tion­ing, the bet­ter, and I think the same is true of per­son­al­ized LLMs—sim­ply next-token pre­dic­tion on raw IRC logs is not as use­ful as being able to in­ter­sperse the LLM’s com­men­tary on what a given state­ment means.

A per­fected GA would be able to do that as it went, but a pro­to­type may have to do a boot­strap of naive train­ing on the orig­i­nal cor­pus, and then prompt­ing/scaf­fold­ing to do use­ful analy­sis to aug­ment the cor­pus, and then re­train­ing, after which it has learned to do aug­men­ta­tion on the fly.

Since we don’t know how to for­mat the data cor­rectly, I ex­pect to it­er­ate. Per­haps we will train naively on just con­cate­nated data, then prompt the GBT for self-analysis and sum­mary, and fig­ure out what kind of an­no­ta­tion is use­ful. Some­thing like, a global prin­ci­pal pro­file PRINCIPAL.md, per-corpus-file sum­maries, and intra-file an­no­ta­tions like <!-- GBT: important: personality -->, and then re-process the orig­i­nal data to aug­ment it, and re­train the foun­da­tion model, and throw in new ‘Q&A logs’ of me with GBT, for ex­am­ple.

Once we fig­ure out good scaf­fold­ing/de­sign pat­terns using our GBT guinea pig, then fu­ture LLMs can just do it out of the box by pre­train­ing on ex­am­ple cor­puses and doc­u­men­ta­tion. And they can just do an au­to­matic ‘im­port’ it­er­a­tion of pre­train­ing, an­no­tat­ing, re-pretraining, and then dy­namic eval­u­a­tion from then on, and let pe­ri­odic model up­grades im­ple­ment a ‘reset’ of re-annotating.

Q&A logs might be ini­tially pop­u­lated using ⁠the in­ter­view prompt: just run it on every sep­a­rate piece of text, and then cu­rate the top 100, say, and sort by in­ter­est. You can ob­jec­tively quan­tify the value of a ques­tion by how many changes it pro­duces to a PRINCIPAL.md or to the meta-annotations in char­ac­ter count or corpus-wide com­pres­sion loss (cf. Gu­rung & La­p­ata 2025)—use­ful sum­maries should make the cor­pus more com­press­ible. Then use that as meta­data and you can prompt to gen­er­ate highly use­ful ques­tions, even­tu­ally.

We can do ⁠ac­tive learn­ing over the log data in the sense of tak­ing a lot of user dat­a­points, doing some sim­ple brainstorm-like think­ing over each one to find the most un­cer­tain ones, and ask­ing ex­plic­itly about the mean­ing.

Then we can ex­per­i­ment with tool-calling (mocked and im­ple­mented by hand) to see how much it cap­tures of my writ­ing process (see my Inkhaven & ⁠Dwarkesh in­ter­view, and ⁠find­ing my ideas), and get a bet­ter idea of what is the best par­a­digm for scal­ing GAs up.

An­other in­ter­est­ing thing to do is to cu­rate self-play data: chat­bot per­son­al­i­ties are well-known to di­verge into strange ‘at­trac­tors’, like the “bliss spi­ral” of Claude chat­bots; so one can char­ac­ter­ize one’s own GA at­trac­tors, and edit data to re­duce them by re­play­ing the con­ver­sa­tion until it starts going off the rails, and writ­ing your own re­sponses.

I do not know which of us has writ­ten this page.

Jorge Luis Borges, “Borges and I”


  1.  

    Iron­i­cally, the biggest suc­cess of the RL sub-field of ‘human pref­er­ence learn­ing’, RLHF, works by not learn­ing any ac­tual human’s pref­er­ence. Those pa­pers which do con­sider the sub­ject, tend to dis­miss the ex­is­tence of human pref­er­ences as noise; for ex­am­ple, the DPO au­thors mea­sure—on sim­ple sum­ma­riza­tion tasks—human dis­agree­ment rates as high as 35% and infer this jus­ti­fies using LLM prox­ies, rather than rais­ing ques­tions about the basic ap­proach of at­tempt­ing to model ‘the’ human pref­er­ence.

  2.  

    Dy­namic eval­u­a­tion is tra­di­tion­ally done as full-model fine­tun­ing, which can ab­sorb ar­bi­trar­ily large datasets.

    For ef­fi­ciency, LoRA can be used ini­tially and is roughly equiv­a­lent to full fine­tun­ing, as shown by Rannen-Triki et al 2024 and ⁠Think­ing Ma­chines 2025, but will even­tu­ally un­der­per­form. (Pos­si­bly LoRAs would be more ef­fec­tive for this pur­pose if done on a very over­pa­ra­me­ter­ized LLM, in which case there may be host­ing ben­e­fits, if a host like ⁠Think­ing Ma­chines Inc. could run many users si­mul­ta­ne­ously to gain through­put.)

    Since we ex­pect prin­ci­pals to gen­er­ate bil­lions of to­kens, be­tween in­gest­ing their archives and syn­thetic data/aug­men­ta­tion and in­ter­ac­tion, it may be pos­si­ble to sat­u­rate LoRAs. If even full fine­tun­ing proves to be in­ad­e­quate, it is pos­si­ble to do model ex­pan­sion by net2net model edit­ing ap­proaches such as adding lay­ers or net­works, like ⁠lad­der net­works, which freeze the orig­i­nal model weights and train a new LLM to ‘use’ the orig­i­nal.

  3.  

    But lifel­og­ging is prob­a­bly not as cru­cial as we used to think, be­cause of­fline data suf­fers from a curse of ex­plo­ration: day to day life is highly pre­dictable and quickly un­in­for­ma­tive about deeper prop­er­ties, be­cause peo­ple usu­ally spend lit­tle time doing un­usual things, and not an­swer­ing strange hy­po­thet­i­cals or in­tro­spect­ing deeply about their pref­er­ences.

  4.  

    Props to the two peo­ple who im­me­di­ately replied, “Be­cause I need a job now” and “I’ll be a nepo baby, my re­tired par­ents al­ready agreed to take me back if I am per­ma­nently un­em­ployed”.

  5.  

    ie. use a ⁠‘rac­ing’ top-k multi-armed ban­dit on­line eval­u­a­tion. My fa­vorite ban­dit al­go­rithm, top-k pos­te­rior sam­pling, is prob­a­bly not ap­pro­pri­ate here be­cause LLM check­points can take a lot of disk space, and will con­tin­u­ally fall out of date and need to be re­trained if they are to be reused later, so hold­ing onto an in­def­i­nitely large pop­u­la­tion of en­sem­ble can­di­dates is highly ex­pen­sive.

Similar Links