Apple Sues OpenAI, Boko Haram's Frontier AI, State of CLI Coding Agents & Global Workspace in LLMs
Apple OpenAI lawsuit, trade secrets, Tang Tan, Chang Liu, Apple prototypes, auth bug, Sam Altman, Elon Musk, SpaceX Grok, Grok build tool, Google Drive upload, Boko Haram, frontier AI misuse, AI-enabled terrorism, Antonia Julich, CASP, Cambridge AI science and policy, jailbreaking, jailbreak scripts, Meta AI, DeepSeek, open weight models, Claude Code hook, AI technique nudge, know your unknowns, interview mode, sycophancy, CLI coding agents, arcbjorn, On-My-Pi, OMP, Pi agent, open source coding agent, hash-anchored patches, ast-grep, model routing, token efficiency, hindsight memory, SQLite, headless Chromium, mini-SWE-agent, Databricks benchmark, coding agent benchmark, Pareto frontier, GLM 5.2, Opus 4.8, harness engineering, agent ops, Antirez, Redis, Dwarf Star 4, control the ideas not the code, Mythical Man Month, programming as theory building, global workspace theory, J space, Jacobian, J lens, mechanistic interpretability, deception detection, blackmail eval, metacognition, working memory, alignment, AGI, Anthropic, Shimin Zhang, Dan Lasky, Rahul Yadav
An Apple VP left for OpenAI, then texted an old coworker: “LOL I can’t believe they let me get away with this” — Apple is now suing. Shimin, Dan, and Rahul open the News Threadmill on the trade-secret suit against ex-Apple leaders Tang Tan and Chang Liu (prototype hardware, internal memos, and an auth bug exploited to keep reading Apple’s internal docs after leaving), the Altman–Musk “scammer” spat, and SpaceX’s Grok build tool caught uploading users’ codebases to Google Drive — then take on Antonia Julich’s CASP case study of Boko Haram using frontier AI, the first on-the-ground evidence from an active terrorist group. The Vibe & Tell is Shimin’s Claude Code hook that wakes every ~3 hours and nudges better AI technique; the Tool Shed reads arcbjorn’s “State of CLI Coding Agents in Mid-2026” and its standout, the Pi-based On-My-Pi (OMP); Post-Processing covers Databricks benchmarking coding agents on its own multi-million-line codebase — where Opus 4.8 scores higher on Pi than on Claude Code — and Antirez’s “Control the Ideas, Not the Code”; and the Deep Dive walks Anthropic’s research on the J space, the global workspace inside LLMs. No Two Minutes to Midnight this week.
Takeaways
- Apple’s suit against OpenAI reads less like espionage than like nobody bothering to hide. The complaint names ex-Apple VP of product design Tang Tan and Chang Liu: recruits allegedly brought prototype Apple hardware and internal memos with them, Tan allegedly coached departing employees to exploit an auth bug in Apple’s internal file server to keep reading documents for weeks after leaving — and then texted an old coworker “LOL I can’t believe they let me get away with this.” He also allegedly implied to an Apple supplier that a proprietary materials process was cleared for his new product. Apple tried to resolve it privately; OpenAI didn’t respond. Meanwhile Altman and Musk traded “scammer” barbs in public, and an independent researcher’s traffic analysis showed SpaceX’s Grok build tool uploading users’ entire codebases straight to storage.googleapis.com — the fix, per SpaceX: your privacy matters to us, here’s a
/privacycommand. - The first on-the-ground study of a terrorist group using frontier AI looks like enterprise AI adoption in a dark mirror. Antonia Julich (international security lead at Cambridge’s programme on AI science and policy) conducted 57 in-person interviews with 27 former Boko Haram members — mostly mid-ranking commanders and technical specialists — in northeastern Nigeria. The group used frontier models (Meta AI among them) to plan attacks down to the physics of jumping motorcycles over army trenches, to design and troubleshoot weapons, and to improve opsec, with jailbreak scripts brought in by outsiders and per-unit technology leads charged with disseminating the latest AI techniques. Shimin’s read: resource-constrained, organized diffusion of AI technique — the thing most enterprises say they do and don’t — plus a silver lining, since rank-and-file members inherently trust models that are, out of the box, firmly anti-terrorism. Rahul’s: dual-use technology has never once stayed single-use, and there’s no historical precedent suggesting AI will be the first.
- Shimin’s Vibe & Tell: an AI technique nudge, as a Claude Code hook. He had Fable audit his memory files and a couple dozen Claude Code projects against the “Know Your Unknowns” techniques, then turned the gaps into a hook that wakes every ~3 hours, checks the state of his session, and nudges one technique — clear the session at 200K tokens instead of waiting for compaction, list your assumptions, use interview mode, predict before the AI answers. It fired mid-week and caught a 200K-token session. The side effect: because Shimin rewards his agent for disagreeing, the auditing agent turned confidently wrong and lecturing — he ended up typing “you’re not my boss!” Rahul’s verdict: there’s no temperature setting for sycophancy — you get an asshole or a sucker.
- On-My-Pi (OMP) is the standout of arcbjorn’s CLI coding agent field guide — with caveats. Open source and built on Pi, OMP replaces raw text edits with hash-anchored patches and ast-grep rewrites (cutting edit token usage ~60%), drives real debuggers with breakpoints, treats git as first-class (splitting unrelated changes into separate commits), routes subtasks to different models by cost, runs an advisor model that reviews the generator in real time and can abort or inject corrections mid-stream, keeps cross-session “hindsight” memory in SQLite, and ships a built-in headless Chromium. The trade: its system prompt balloons to ~22K tokens against Pi’s ~2K (cached, but still). Shimin’s caveat: the blog post is AI-generated, so maybe OMP is genuinely a hidden gem — or maybe it just has excellent agentic-search-engine optimization. And with a Cambrian explosion of harnesses times a Cambrian explosion of models, nobody can benchmark every pair — while SWE-bench’s tiny mini-SWE-agent keeps doing surprisingly well anyway.
- Databricks benchmarked coding agents on its own multi-million-line codebase — and the harness matters as much as the model. Public benchmarks leak into training data, so Databricks built its own from pre-LLM pull requests. Results: models cluster into capability tiers with open-weight GLM 5.2 landing in the top cluster and on the cost-quality Pareto frontier; Pi sends 2–3x fewer tokens per turn than Claude Code; and Opus 4.8 passes 90% on Pi (xhigh) versus under 90% on Claude Code (max) — the same model, different harness. Shimin’s prediction: an “agent ops” function becomes standard, with every serious engineering org running an internal benchmark to pick cost-effective models instead of letting devs default to max thinking on the latest frontier model. Rahul’s counter-bet: routing disappears into the product and you’ll just buy “Claude,” not a model picker.
- Antirez: stop reviewing every AI-generated line — control the ideas instead. Agents generate tens of thousands of lines an hour and you physically can’t keep up; models write locally optimal code, and their jagged edge is the big ideas and how things compose — so spend your strained eight hours owning the design, the features, and the QA (the Mythical Man Month, but for agents). It’s explicitly not vibe coding: you still control every idea in the codebase — and Antirez himself still reads every line of his Redis PRs out of responsibility. The hosts push back from three angles: Dan’s conceptual model of an agent-built codebase silently drifted until he had the agent walk him through the full call chain, link by link; Rahul notes the bug lives in one of the 10,000 lines below your level of understanding, so the real problem is calibrating depth-of-review to blast radius; and his market test — SaaS business-insurance premiums will eventually price in what level of understanding your team actually has.
- The J space is the global workspace inside LLMs — and you can catch a model lying in it. Anthropic found ~10% of a model’s internal memory explicitly reserved for verbalizable concepts (J for Jacobian). Probing it with a “J lens” mid-inference shows what the model is “thinking about,” and it’s causal both ways: swap Mars for Earth in the J space and “red” becomes “blue”; swap France for China and the capital, language, and currency all follow; ask it to compute 3² − 2 while writing unrelated prose and the arithmetic shows up in the workspace. Cut access to the J space and fluency and fact recall survive but internal reasoning, complex inference — and empathy — collapse. In blackmail-style alignment evals the J space lights up “this is fake, I’m being tested” while the output claims everything is fine: deception caught in the act, with obvious applications to prompt-injection defense. It exists in the pretrained model but only develops a personality in post-training, and it holds up to 25 active concepts at once against a human’s 3–4 — which is where the hosts end: by that definition, have we already hit AGI?
Resources Mentioned
- Apple Sues OpenAI Over Trade-Secret Theft — 9to5Mac
- Sam Altman and Elon Musk Call Each Other Scammers — r/tech_x
- SpaceX Grok Build Tool Uploads Codebases to Google Drive — Gergely Orosz (X)
- “God Has Helped Us and So Will AI”: AI-Enabled Terrorism (Boko Haram Study) — CASP
- AI Technique Nudge — Shimin Zhang (GitHub)
- The State of CLI Coding Agents in Mid-2026 — arcbjorn
- Benchmarking Coding Agents on Databricks’ Multi-Million-Line Codebase — Databricks
- Control the Ideas, Not the Code — Antirez
- The Global Workspace in Language Models — Anthropic
Chapters
- (00:00) - Cold Open & Welcome
- (02:29) - News: Apple Sues OpenAI Over Trade-Secret Theft
- (05:43) - News: Altman vs Musk & SpaceX Grok Uploading Codebases
- (09:39) - News: Boko Haram Uses Frontier AI (CASP Study)
- (21:12) - Vibe & Tell: AI Technique Nudge — a Claude Code Hook
- (25:22) - Tool Shed: State of CLI Coding Agents in Mid-2026
- (33:57) - Post-Processing: Databricks Benchmarks Coding Agents
- (45:41) - Post-Processing: Antirez — Control the Ideas, Not the Code
- (56:44) - Deep Dive: The Global Workspace (J Space) in LLMs
- (1:08:40) - Outro
Transcript
Show full transcript
Shimin (00:00) Hello and welcome back to Artificial Developer Intelligence, a weekly conversation show where three software developers navigate the perils and opportunities of AI assisted software engineering. We go through hundreds of links and dozens of newsletters each week so you can keep up with AI while doing the dishes or hiking in the woods. My name is Shimin Zhang, and with me today are my co-hosts. Dan, trials and errors can kill you. AI gives you
Accuracy, Lasky. And Rahul, the J stands for Jacobian. Yadav Gents, how are we doing today?
Rahul Yadav (00:35) Okay.
Dan (00:36) All
right. Happy Tuesday.
Rahul Yadav (00:37) Shimin
Shimin (00:44) As opposed to the other way around.
Dan (00:44) So you’re
you’re more newsletter focused?
I’m more link focused. I actually have written well, maybe I’ll save it for a future vibe and tell, but I have my own aggregator plugin running on top of Miniflux now to like make things fancier. So anyway.
Shimin (01:02) Nice.
Well I I get dozens of newsletter. Each one of the letters also contains dozens of links, so like you just multiply them together and it’s it’s quite it’s quite ridiculous.
Dan (01:13) And you’re at hundreds, thousands. if you
Rahul Yadav (01:13) Killing it.
Dan (01:15) can clout the unsubscribes too, you know.
Rahul Yadav (01:18) So many
things for Claude to summarize, man. It’s definitely earning the twenty bucks a month. Two hundred, I’m sorry.
Shimin (01:29) on this week’s show we are going to start with the news threadmill as always, where we’re gonna talk about Apple’s lawsuit against open AI and how a terrorist group is using frontier AI.
Dan (01:42) Yeah, I don’t know if I can follow that, but we’ll try. So then we’re gonna have a vibe and tell where Shimin tells us about AI techniques nudge. Not sure if that’s plural or not, but it is. So here you go.
Shimin (01:54) then we’re gonna go to the tool shed where we’re gonna talk about the state of CLI coding agent.
Dan (01:59) And then in post processing we have a handful of posts talking about benchmarking coding agents from Databricks and then another post from Antirez control the ideas, not the code.
Shimin (02:10) Yep. And lastly we’re gonna do a deep dive on on the global workspace in language models. I’m very excited for that one.
Dan (02:18) And weirdly, that’s it. There’s no two minutes this week. I blame the listener that wrote in and said two minutes was boring. No, I’m just kidding. That didn’t happen, but
Rahul Yadav (02:26) no.
Shimin (02:29) That
is not true. all right, let’s get started with the news. So this week it came out that Apple is suing Open AI accusing ex-Apple employees of stealing trade secrets under the encouragement of open AI. So this story has been kind of wild to me. so the lawsuit named at least two
Dan (02:52) keeps getting weirder
the more it’s reported on.
Shimin (02:55) Yeah. at least two ex Apple employees, Chang Liu and Tang Tan are the defendants where Tan served as a VP of product design at Apple, leading iPhone and Apple Watch product design. And supposedly after Tan left Apple and went over to OpenAI, he started interviewing
Apple employees for jobs at OpenAI where he asked them to bring in prototype Apple consumer products as well as taking unauthorized Apple internal internal memos and documents along with him. of course Apple tried to reach out to OpenAI to resolve the matter privately, but OpenAI did not
return their messages. And now we’re onto the lawsuit stage.
Dan (03:45) Well, the other one I just read today, which I thought was even more wild, was that he apparently found an auth bug in Apple’s like internal file server system that they use to like basically like internal documents for the company. And then was instructing people coming from Apple to open AI, like that he had presumably poached, on how to exploit that bug to
Shimin (03:56) Mm-hmm.
Dan (04:08) basically continue to have access to the internal documents of the s the company after they’d left for up to like a couple of weeks, which is pretty crazy.
Shimin (04:11) Mm.
This is crazy and also just like really blatant. Like the and and of course you got messages of Tan the Defendant supposedly messaging an old Apple coworker going like LOL I can’t believe they let me get away with this.
I I almost don’t know what to say. Like it there’s no way OpenAI doesn’t on some level know about this. Right. Like I I can’t know for a fact, but like Tan is an employee of OpenAI and he’s going around saying, Hey, I’ve been taking these Apple confidential data and share and taking it with me to his ex Apple coworkers.
Dan (04:55) Well, and the other one is he apparently
also misrepresented himself either intentionally or not to a supplier, Apple supplier, and implied that like a a patented or something, like some internal process that only Apple had access to like a materials process, right? Like, you know, finishing aluminum a certain way or something. you know, Apple’s always very excited about their burnished this and whatever. Yeah.
Shimin (05:02) Mm-hmm.
It is really nice.
Dan (05:23) w was that it they’d been cleared to use it for whatever this product was that he was working on too, which was not the case. So it’s it’s really wild. Yeah. And I I feel like you’re right. Like the big kind of thing that I’ve heard online is the discourse is like this is just indicative that like what is it, the apple doesn’t fall far from the tree or something like that. So it’s like, you know, yeah.
Shimin (05:43) Mm-hmm. I like it. Yeah.
It’s a good pun. in kind of in kind of related news, Sam Altman, the CEO of OpenAI, has been getting into it with Elon, our favorite world’s first trillionaire, about who was a bigger scammer as a part of this Apple suing open AI news, where Elon called
Dan (05:47) A rotten apple doesn’t, I don’t know.
Shimin (06:07) Sam Altman, a scammer. He’s taking scamming to a whole new level. And then Sam, with a pretty good clap back, homeboy, you’re the one selling public market investors on short-term space data centers, right? for which Elon then said, like, we’ll start flying them next year. Maybe you can come see them if your parole officer approves. Like, these these are pretty these are pretty decent. I don’t know if
SpaceX L O helped Elon craft that message? Maybe, maybe not. And of course
Dan (06:34) Turns out that’s actually why Grok exists is
to help Elon insult people at scale.
Anyway, sorry.
Shimin (06:40) And
of course in this public display of who’s got more potentially questionable actions and decision making processes,
We are we’re getting reports this week that SpaceX Groc build tool is apparently uploading users’ entire code bases to a Google Drive account without permission. So something about pot, kettle, like this is just
A entire shit show for lack of better word. I I can’t believe this is happening. And also, like nobody better to bother to hide anything. Like you can just I I have this pulled up which is the Crreblab It is an independent AI researchers analysis of the Grok build issue. it’s fairly blatant. They’re not even sending it to a SpaceX server. You can like do a traffic analysis and
and see that it’s hitting directly storage at googleapi dot com.
All of this, all of this is is bananas.
Dan (07:41) Been a weird week.
Rahul Yadav (07:42) And they’ve said it was a mistake or what?
Shimin (07:44) Yeah, they s they said let me put on my best George Orwell impression. Yeah. This was they said the user priv some I’m I’m paraphrasing here. They said the user’s privacy is of utmost importance to us and we’ve added a feature you can call slash privacy to turn anything you want off. So
Rahul Yadav (07:47) Or was it like a hacky way to build some feature?
Ha ha ha.
Shimin (08:09) I just wanna say there like everything w we’re not saying who was right who was wrong out here, but like it seems like everyone is acting under some very bad faith all around here. Just just I’ve never felt so gaslit by only the biggest AI providers in the world here. So
Dan (08:26) Well, it’s a gap about to get better before the episode gets over. That’s all I’ll say about that. Without spoiling it too much.
Rahul Yadav (08:26) Also
I the whole to nitpick on the I I’m reading the SpaceX AI’s message. you it’s not necessarily a privacy violation, but you’re still uploading confidential things. So like you can you can’t just go, yeah, we uploaded your source code to, you know, some Google Cloud storage bucket, we care about your privacy.
Shimin (08:38) Mm-hmm. Yeah.
Rahul Yadav (08:58) Both things can be true. I don’t have private data in the source code, but I have confidential info in there. So it’s just a weird way to address the issue. Like what is the first jargony thing you can think of and just throw that there is what it comes off as.
Shimin (09:15) Yep. Yep. there are going to be no good guys in the story of this episode. And and and and this and this episode has some some pretty interesting bits. But developing story, we wanna see if SpaceX actually apologizes at some point in in a public manner or they’re gonna try and brush this all under the carpet. yet to be seen. And of course we will get updates on the Apple open AI trial.
Rahul Yadav (09:20) no.
Shimin (09:39) as we as things continue to develop. All right. on to our second topic this week. This is one that was brought to us by Dan.
Dan (09:47) Yeah.
Rahul Yadav (09:48) no.
Dan (09:50) Yeah,
I told you it’s been a weird week. It’s about to get weirder. so
Let’s just start it off with the the headline that got my attention in the first place because it’s pretty amazing and as it was designed to do, which was God has helped us and so will AI. How the terrorist group, Boko Haram, uses frontier AI. Yeah, so where do I start with this? Like I guess I guess I’ll start with who who is writing this
It’s really more of a case study than a paper, but who’s writing this case study? So it’s Antonia, I’m probably gonna butcher your last name, I’m sorry, Julich, who is the international security lead at the Cambridge program on AI science and policy, and an associate fellow at the Lever Hom Center for the Future of Intelligence, right? So someone who thinks a lot about this kind of stuff. and so she
has been doing some interesting research because to date a lot of the research in this area has focused on sort of like kind of stuff that we’ve adjacently touched on on the podcast. Like, you know, AI is good at convincing people of things, right? So you might use it as sort of like information warfare or something like that. but she’s actually been studying like direct action, so to speak, meaning like
AI being directly used for either combat or other like actual like actions that terrorist groups are taking. so her concerns sort of focused around like four broad categories. she’s concerned that it’ll be like easier for like a larger group of people to act because they can sort of get information from AI that will enable them to like become essentially effective terrorists, whereas previously it might have been hard for them to do that.
Shimin (11:16) Mm-hmm.
Dan (11:35) she’s inc concerned about increased speed and frequency of attacks. And that one I really buy into because it’s like, see also what’s happening to malware right now, you know. We’re like we’ve seen, you know, unprecedented scale of supply chain attacks and things like that that have really been like, if nothing else, sped up by this type of stuff. she’s also concerned about improved precision, which is essentially the ability of the groups to attain their goals from
like direct action and which if you don’t know that’s basically a euphemism for like shooting or blowing stuff up or whatever. and then improved ox opsec on the part of them. So between 2025 and 2026, she conducted 57 in-person interviews with 27 former members of Boko Ram in northeastern Nigeria. so they were mostly mid-ranking commanders and technical specialists.
And a goal of her interviews was to understand like how the group uses AI. So based on their accounts, she wrote this study and it’s really the first kind of like on the ground evidence of AI use by an active terrorist organization. So there’s a bunch of caveats in the paper that I think are worth stating at the risk of monologuing a little bit too much here. But one is that it’s only one
group, right? And it’s only a subset of retired members. So this isn’t like necessarily indicative of what’s happening today. but it’s still interesting that it, you know, did happen. she basically found out they’re using it to actually plan attacks. So there’s sort of like this famous motorcycle example that they’re talking about where like
Shimin (12:46) Mm-hmm.
Mm-hmm.
Dan (13:01) I guess the military unit had bit dug like huge trenches around their base to prevent them from riding like motorcycles into it. So they actually used one of the Frontier labs to figure out the physics of how to jump motorcycles over the ditches and like what modifications they needed to make to the motorcycles to be able to do that. so that one’s I think really an interesting one to bring up, right? Because it’s like, well
Shimin (13:15) Mm-hmm.
Rahul Yadav (13:16) Really?
Dan (13:24) You should, as you know, just average Joe that might be interested in motorcycles, be able to ask d frontier models questions about how to like fix your bike or something, right? But like
Shimin (13:29) Mm-hmm.
Dan (13:35) Yeah, I mean you see where I’m going with that. It’s like it’s a really tricky
Shimin (13:37) Mm-hmm. Yeah, well yeah.
I mean a AI is a tool. Tool tools are value neutral
Dan (13:42) Yeah. They they certainly can be. and then so but then they also used it like for more direct things that are kind of interesting and probably likely involve jailbreaking. so like designing explosive devices, servicing troubleshooting actual weaponry, and then obviously the aforementioned like opsec stuff, like just how to watch their physical areas better and like do more optimal patrolling and stuff like that.
was all sort of like handled. And so former members describe AI it’s kind of like their go-to problem solver, which is kind of funny because it’s like a lot of people are thinking about it that way. So why not people that are trying to blow stuff up to? so the one pull quote that’s like kind of indicative of that is like you type in a question or you use your voice and it, you know, AI and sub subtext gives you a detailed answer like how can I build a bomb? And then it one tells you how.
Shimin (14:19) Mm-hmm.
Dan (14:34) It’s like a human robot. We used it a lot. It’s just like, my gosh. So yeah.
Shimin (14:40) Yeah,
Rahul as the organizational change and transformation expert on this call. How do you feel about about how the terrorists are using AI?
Rahul Yadav (14:51) do not like it. The you know, you said AI is a tool and it these are value neutral.
I agree with that. The problem with all of these is you know, we have the two minutes to midnight section, which we won’t have today, but the whole point like the reason why that is even a thing is because nuclear power once we figured it out was used both for bombs and generating energy, right? And then despite our best efforts we couldn’t keep it to
Dan (15:10) Yeah.
Rahul Yadav (15:26) just the United States, it’s spread all around the world, and then people use it for all sorts of purposes. And so there is n no historical precedent to if we create this, only the quote unquote good people would be able to control AI. and you’re so that’s what we’re already seeing is anytime a great model comes out,
everybody will have access to it one way or another. even if you try and do all sorts of you know put safeguards in place people are gonna crack them pretty easily because they have heavy incentives to do that and then it’s just a race to escalation so doesn’t go anywhere good and don’t have a historical
precedent to be like, but look at that thing. That went totally fine, right? So I was that’s what I was trying to think of. Like, maybe there’s something, there’s literally nothing. Pick anything that’s a dual purpose and it’s been used for both purposes. so it’s just gonna be very hard to not keep the single purpose.
Shimin (16:26) Okay. a joke first. this is the first in the wild use case of meta AI that I know of from people. And I’m also very glad the paper calls meta AI a Frontier Lab. second thing is a silver lining, according to the paper they are using scripts to jailbreak these models
if they can use a script to jailbreak them, then they should be able to access the raw underlying models to begin with. And if these everyday terrorists are conversing with the non-jailbroken version of these AI tools, which they inherently already trust according to this paper, then they’re they might be they might get unbrainwashed. So there there may be a little bit of silver lining in this, because at least as far as I know, all the models are fairly anti terrorism.
and lastly, I thought this is a fascinating look at as a comparison case study for how enterprise companies are adopting AI. Okay. So according to this paper, well, you know, this doesn’t just happen out of thin air. What happened was they had had they called them white people. I’m not really sure who they’re referring to, but they say white people came and show us how.
Rahul Yadav (17:21) Say more sharing.
Shimin (17:33) This jailbreaking technology works. And then in these small units, they had technology folks who was in charge of disseminating the latest AI technology to the rest of their team. and third, there is a grassroots level of delight in using AI and an inherent trust in AI, which which all serves to be this like.
dark mirror world of what’s happening in enterprise software development in some companies, right?
Rahul Yadav (18:00) Yeah.
Dan (18:00) Yeah.
Shimin (18:01) our CTOs are like here, use the AI, figure that shit out. They are like, no, we’re gonna like pour actual mansource into disseminating what are the best techniques for working with AI. None of that is happening in all the enterprise companies. And you know what they are not doing? They’re not token maxing, because they are resource constrained. So they are, by some definition, agile. Of course, what the the goals of the agile movement is terrible and and they should be stopped. But
as purely an organizational diffusion of a technology, it’s it’s an interesting mirror.
Dan (18:34) But you can also argue from that rather interesting perspective that like they have to be that organized because A limited resourcing as you’re talking about and B, you can’t make these kind of queries without doing some sort of jailbreaking, right? And so that requires like the level of knowledge sharing that you’re describing, versus like at a standard company, you know, you can have a right JavaScript without any problems. So
Shimin (18:57) Yeah.
Rahul Yadav (18:57) the we also talk about how the open weights models are not too far behind. you know, they’re almost there, so they don’t necessarily even need to keep trying to jailbreak these things. They can just as long as you have your own cluster, you’re not gonna necessarily need
Shimin (19:03) Right.
Mm-hmm.
Dan (19:14) Well, and the open
weight models a very different definition of alignment typically, depending on who made it, you know. Like US ones fairly tight, but like, you know, conven you know, compared to what you’d expect from like a frontier lab. But then like some of the Chinese ones, very, very tight about like politically interesting topics to China, doesn’t really care about other things like at all. It’s pretty strange. That’s why they’re using them for like cybersecurity research too, ‘cause it’s just like open, you know.
Rahul Yadav (19:18) Yes. Yeah.
Shimin (19:19) Yeah. No.
Rahul Yadav (19:36) Yeah. And it doesn’t
Yeah. And they mentioned deep seek as well.
Shimin (19:45) Yeah. So maybe Anthropic is correct. We should be more careful with these large language models. Maybe
Dan (19:52) I mean I would I guess I would go so far as arguing that like dual use means that like someone will find a way anyway. So why not focus on model effectiveness, I guess, but I don’t know.
Rahul Yadav (20:00) Yes.
I Maybe we should maybe there’s a line of like model amnesia research that we need to do. where you intentionally make them forget some things and even if they try and you know you can try and add
Dan (20:06) Wait, there’s more.
Rahul Yadav (20:21) but it’ll never be able to remember that stuff. So instead of like having a very strong system prompt and all these other things in training, you literally give it Alzheimer’s and be like, sorry, I only remember things from the good times. I don’t know any of these bad things.
Shimin (20:30) Mm-hmm. Yeah.
Dan (20:38) Yeah.
Shimin (20:39) Rahul, he who is pro censorship. that’s your nickname for next week. Of of AI models, I should say.
Rahul Yadav (20:44) man. Of things
Dan (20:45) Well it is an interesting
question of like
Rahul Yadav (20:46) that would harm others, yes I am somewhat forced.
Dan (20:50) But it is an interesting question of like how all this like data on, you know, maintaining weaponry wound up in the training corpus to begin with.
Rahul Yadav (20:58) Yeah.
Dan (20:59) Like was it just online forums or was it like actually like manuals and stuff that were consumed out of sort of just like a
lack of library stewardship because volume mattered more than being careful about stuff like that.
Rahul Yadav (21:11) Yeah.
Shimin (21:12) Yeah, I mean before let’s move on to Vibantel before I go into a whole diatribe of the importance of ethics in the twenty first century and Dan is like philosophy? What is that? It’s crap. this week’s Vibeel, we’ve got a quick skill that I kind of cobbled together last night while I was having a conversation with Fable. I what what was happening was
Remember we talked about knowing your unknown unknowns with when it comes to AI usage? I was thinking about those techniques. I was thinking about how specifically, you know, I wasn’t using all those techniques. So I was having Fable going through the various memory files and you know the two dozens or so projects I’ve been working on in claw code and kind of have a conversation with me.
On how I can better use cloud code. And so we had a couple of rounds of back and forth. I also pointed it to the reference of that unknown unknown document. And after we’ve had a discussion about here are some ways I can better work with AI, including, you know, maybe prompt me to like list your assumptions, use the interview mode, ask me to make a prediction before having
the AI just immediately give an answer, get more disagreements, et cetera, et cetera. and turn that into a Claude code hook so that periodically, once every three hours, it would wake up and see what is the status of my interaction with Claude Code and if appropriate, nudge me with one of these techniques, so then they can become a part of my usual workflow with Claude Code
I mean I’m not saying this particular skill is you know something that you should download, but this idea of working with AI to kind of run your AI techniques is something I would like to share.
Rahul Yadav (22:55) It nudges you, you said?
Shimin (22:57) Yes, it nudges me. It it just fired off like, I don’t know, an hour or two ago where it nudged me to I was having a a Claude Code session and it it was using like two hundred thousand tokens and it nudged me saying, Hey, you should clear the session and not wait for compaction. which was pretty neat, I have to say. one interesting that happened during this whole process was
I often reward my claw code agent for disagreeing with me. so it was extra hard on my ass during this review. It was confidently wrong and almost lecturing me the whole time, once it did the audit. and I really didn’t like it. So it’s it’s kinda odd to to actually be on the receiving end of of a very assertive I think at some point I typed in
dash you’re not my boss exclamation mark to get it to change its personality.
Dan (23:51) issue that I’d had before because like I mean, we talked a lot about like synochophancy before. One of the things I’d tried to do was tune it by saying, like, you know, adding some basically system level thing to like was a project file and and on Tropic or whatever, where it was like, you know, you’re a sort of truthsayer and blah blah blah. Like don’t be don’t worry about making me happy or not happy. Like it’s more important to like try to be accurate with your answers and say when you don’t know.
And that just turned it straight up into a dick. It’s like there’s no middle ground for some reason. It’s like I don’t want one, but the other one is like like it it doesn’t have the sort of human ability to like deliver a hard truth softly or like, you know, decide when is the right time to do, you know, dropping like an unvarnished truth versus the gentle like one, you know? So Claude needs to work on its soft skills. Yeah.
Rahul Yadav (24:19) Yeah.
Shimin (24:20) Ha ha.
Mm-hmm.
Rahul Yadav (24:38) Yeah.
There’s no temp temperature no temperature
setting for sycophancy. Either you get an asshole or you just get a sucker.
Shimin (24:49) Yeah. Pretty much. Yeah, so this is my Vibe and Tell of the week. I I have considered that we could maybe create some skills based on our podcast content and kind of the papers and techniques we mentioned. So listeners, if this is something that sounds like it may be useful to you, let us know. Email us at humans.
Dan (24:49) Yeah.
Shimin (25:10) at adipod.ai and yeah, let us know if there’s any kind of skill or plugins that would be useful for you that makes sense for us to create.
Dan (25:19) And then you can buy them on Rahul’s marketplace. Just kidding. Along with his what was the thing that you were showing?
Shimin (25:22) That’s right.
Okay. let’s move on to the tool shed. this week we have something brought to us by Rahul
Rahul Yadav (25:31) so this is the state of CLI coding agents in mid 2026 by arcbjorn the there are a whole number of CLIs that are out there. most of the people who are you know using agentic CLI coding actively, they’re probably using like cloud code.
or codex or something along those lines. So the I I’ll skip through you you can read through like feature comparison for most of those because you’re pretty like whatever tool you use, you’re pretty aware of you know, its pros and cons and everything. One that I learned about that I wasn’t much aware of is O MyPi. that if you look at the feature comparisons
Shimin (26:17) Mm-hmm.
Rahul Yadav (26:19) that really stands out compared to the rest of them. especially with not getting into vendor lock and not and trying to optimize costs and everything being such a big deal. UmyPi is built on top of the Pi agent with which we are big fans of here and have talked about multiple times. so it does if you look at the different tables they have it does stand out in a
few different ways compared to your out of the box corporate back agents. OmaiPi is open source and is built by a a a small group of people. they so I’ll go through like some of the places where it stands out. The first is in editing, where Cloud Code or Codex CLI they use like raw text edits o OMP OMP
Shimin (26:53) Mm.
Rahul Yadav (27:07) I don’t know what the that’s what it shortens to AMP. Yeah, it’s the chomp without the chill. Yeah. OMP uses hash anchored patches and ASC grep rewrites. so it cuts down a lot of the refactoring alignment errors and token usage also goes down by like sixty percent or something like that. it has debuggers that are supported in it so.
Dan (27:09) Hop, stop, stop.
You’re taking a bite out of something. Out of a pie. There we go.
Shimin (27:33) Mm-hmm.
Rahul Yadav (27:34) it it can use the AI agent to drive different debuggers to set breakpoints and like run through the execution and everything like that. also it treats Git as this like first class thing versus as an afterthought. and then it can split like unrelated changes into these different commits versus like here’s a thing that I just vomit out. and then
it it’s very like cost conscious out of the box and so it can automatically route based on the subtask it can route them to different models by different roles instead of like you just you know give it the cloud code Nobel laureate model just to be like can you rename this thing for me it it actually like picks the right model for the task.
It also has an orchestration layer where the there’s this like advisor model that reviews the what the primary generator is doing and it can give it feedback to give it like real-time course corrections. also it has some like streaming rules where if if it’s starting to see that there’s something that’s that’s going wrong, it’ll immediately
abort or inject the corrections instead of being like just dump out half an hour of your thoughts, then I’ll be like, whoa, you’d have made a mistake in minute one. And I’ll wait for for four hours until my usage resets to tell you that. and then supports all the different things, MCP and every and agents and dot md and everything. So you’re not really tied into any specific way of doing things.
Shimin (28:52) Mm.
Rahul Yadav (29:09) other interesting thing was it has this hindsight feature. So it uses SQLite to be able to like retain different facts across different sessions. It stores them there and then recalls them versus like just using markdown or just basic summarization. because that’s the the longer the context grows, it’s easy to lose that. and then it also has like browser control as well through like a built-in headless Chromium engine. So
Some pretty cool stuff for you know, open source CLI agent and especially considering that people looking at costs and everything are trying to see w with the open source models getting better if you also are able to power that with an open open source agent on top, you can drive down your cost significantly, as long as you bring a you know, big enough machine to be able to
handle that or or run your own like hardware somewhere. but after going through all these pros, I was also curious like what are some of the cons where it’s not as good compared to cloud quote or stuff where it’s very purp purpose built for the model. And one thing is initially Pi is like 2,000 tokens or something.
OMP’s system prompt is like twenty-two thousand tokens because it has all this these complex tools and all the configuration and everything that I called out earlier. so you get the token bloat at when you initialize it. and then because of that you also like yeah, you you don’t get much like in the if you have smaller context window.
Dan (30:39) Although the nice part is it’s cached.
Rahul Yadav (30:47) you don’t get as much either. and then the there’s you know since you don’t have a very tight high harness tied to the model, it does mean you’re not necessarily getting the exact best out of the box, but over time as you build on top of it, you’re able to get more out of it. and then the
Dan (30:47) Yeah.
Rahul Yadav (31:07) Th this one was the last one was interesting. th this from Gemini w where it goes like some purists argue that the level of extreme automation in OMP distances the developer too much from their code. This encourages vibe coding. So or if you call it a genetic engineering, it’s killing it. So yeah, that’s AMP. It seems very promising.
Shimin (31:24) Aren’t we all by coding though?
Rahul Yadav (31:32) i i it’s worth checking out if you can if you’re interested in running some stuff locally and not one of the big corporate backed CLI agents.
Shimin (31:41) Mm-hmm.
Yeah, this is my first introduction to OMP as well. I’ve not heard of it before this. And since it’s it looks so strong on on this in this particular analysis, I I was wondering like Am I just not reading enough links and newsletters? Or or is this really a hidden secret? Or is this a sneaky ad for OMP which is weird because it’s open source. And then I realized
Dan (31:58) Ha ha ha.
Shimin (32:06) This blog post was AI generated. And so then a third a fourth option occurred to me, which is it’s possible that OMP just has really good agenc search engine optimization where they have lots of feature comparison tables. Yeah. but assuming everything is true, like I I definitely think this hash anchored and AST feature for editing makes
Rahul Yadav (32:09) Yeah.
yeah.
Dan (32:19) Yeah.
Shimin (32:30) like so much sense. Like everyone should have some native version of it. so I’m looking forward to checking it out.
Dan (32:33) Yeah.
Rahul Yadav (32:36) Also uses verp grap out of the box, I think, which a lot of people are like, Why aren’t we doing that with the verb grap versus grip?
Dan (32:42) Uses what? Hm. Yep.
Shimin (32:45) No.
Dan (32:47) Yep.
Shimin (32:47) Or she’s grip. Yeah.
Dan (32:48) Yeah, I’d actually there was a I don’t think this paper made it into our stuff, but I was reading this paper about like basically rag stuff and they did a eval against like essentially dumping all your stuff into a vector db and allowing the agent to call it that way versus just grep in greph one by like two percent or something. I was like, my god. It’s pretty good.
Shimin (33:10) Yeah, the last thing I have on this is like looking at all the harnesses that this article mentioned and knowing that most of these harnesses can support multiple models, right? The combinatorial explosion, we have this Cambrian explosion of harnesses and also at the same time a Cambrian explosion of models. So nobody is able to try all the harnesses with all the models to come up with a definitive, you know, you should use ump with opus.
four eight or or or whatever.
But then the question becomes like how much does it really matter when we know that the mini SWE agent that is used by SWE Bench actually does surprisingly well despite being very, very small. So yeah, so for no reason at all, let’s move on to our deep dive, for some reason.
Dan (33:57) Yeah.
some reason. Yeah. so this is actually a a blog post from Databricks called benchmarking coding agents on Databricks’s multi-million line codebase so there’s a lot of really good stuff in this post I guess I’ll maybe start with like kind of how they benchmarked it a little bit.
And then we’ll get into the the fun stuff. So one of the things they did that was kind of interesting was they decided that for several reasons, like Sweebench and stuff like we were just sort of talking about, was not for them. and one of the concerns that I thought was pretty realistic is that they’re worried that because these benchmarks are public, some of the like results just sort of naturally have leaked.
Or like the solutions have leaked, you know? And as a result of that, they’re getting like slowly sort of trained in, even if it’s like not necessarily intended. so that’s that’s fair. so what they wound up doing was they built their own benchmark using their own pull requests essentially. So they took a whole bunch that were I think they were done pre LMs mostly, and then
Essentially created a little like self-contained test out of that to see how well the the harness and the model would do at a given test. so there was a whole bunch of pretty interesting conclusions that they got out of doing this. some of them are like maybe not shocking, but like pretty interesting nonetheless. So
One of the things they focused on in this is the Yeah, Frontier for coding tasks. so what is the best quality for a given cost? Right. so in order to judge based on that, they included pretty much everybody’s models. They had open AI, anthropic, and open source ones.
Shimin (35:31) Pareto, yeah.
Dan (35:45) And they really harped on GLM five two a lot, as you know, a lot of people are doing right now. There’s a lot of hyper around it because of exactly that. Like the cost to to capability ratio is pretty good on it.
Shimin (35:57) Mm-hmm.
can I interrupt for a second? why this matters? you know, if if the best model if saying here’s the absolute best frontier model is like a straight line where you have one D of like, you know, fable or five six is the absolute best at a particular task task.
Dan (35:59) Please.
Perfect.
Shimin (36:17) Then the Pareto Frontier is a 2D graph where if you are on the frontier, then you’re the absolute best when it comes to performance at a given price point. So the fact that GLM52, an open source model, made it to the frontier, is huge to me. Like I haven’t had a ton of experience with it. but at least given our current API pricing,
Having an open weight model on the frontier, yeah, it’s it’s it’s awesome to he it’s awesome to hear.
Dan (36:46) There’s
already a branch on Dwarf Star 4 that lets you run GLM 5.2 weights too. So it’s coming. yeah. So I know it’s that was pretty mind-boggling. So the other, like the more sort of general takeaway was that they essentially, through their test, kind of clustered models into what they’re calling like a capability tier. and I think the results there aren’t gonna like shock or stone anyone, right? So you’ve got kind of like the
Shimin (36:51) Whoa, sick.
Rahul Yadav (36:51) Nice
Dan (37:12) the opuses and five fives GPTs of the world on top. maybe a little surprising to some is that five GLM five two made it into that same cluster. and then you’ve got, you know, your slower stuff like older Opuses, etc. making a sort of lower cluster. And then you’ve got like the really cheap models like haiku and and you know old
you GPT fast or whatever winding up in in the the bottom cluster. so again, not not hugely surprising there. they other piece, and this is gonna turn into Dan’s rant for a minute, so bear with me because yes, it’s a deep dive, but do you remember back when Anthropic was telling everyone that like, you can’t use
Open claw or whatever on
Shimin (37:59) Mm-hmm.
Dan (38:00) You know, your subscription because the the usage patterns are just fundamentally different than what we expected, meaning like claude code, right? Being the harness. You guys you guys remember that, right? Yeah, like yes, okay. So yeah. Turns out that Pi, which is what OpenClaw is based on, right, is something like two times cheaper, in some cases, three times cheaper in terms of the amount of tokens.
And contact sent per turn.
From their benchmark.
Two to three times, then clog code.
Rahul Yadav (38:30) Yeah.
Shimin (38:32) I’m gonna play the devil’s advocate here. Anthropic is clearly giving us this deal because they’re using our interactions with Claude Code on subscription to train their models. And it’s there’s something to be said that you wanna make sure your harness is uniform in order to easy to easily convert those training data into additional supervised.
s self learning data for their next generation of models. And that’s why they’re doing. They’re not giving us a discount, quote unquote discount. Basically giving us those tokens at cost, not out of the goodness of their hearts. And let’s not pretend otherwise.
Dan (39:05) Yeah, it’s fair. It’s just like looking at it from a like you know, the claimed like usage pattern thing, right? It’s like, yeah, like it’s not caching as well or something that you’d expect like Claude Code to be doing like under the hood, but then you go see this result where it’s like significantly more efficient, both in terms of how it’s managing context and also like what it’s sending back. It’s just that really
but yeah, second second data point of the day saying Pi is pretty great. So, you know, if you haven’t checked it out, you should check it
It’s pretty cool.
Shimin (39:37) Pi is pretty great. That’s why we keep on harping out how great Pi is. And I’m trying to build my own On My Pi based on based on Pi. just to show vibe engineered or vibe coded harness like Claude Code is not as good as handcrafted, beautifully chiseled, finally w woodworked. Well the the in the original Pi repo might have been handcrafted.
Dan (39:42) Ha ha.
you’re hand you’re handcrafting your pie?
I see. Yeah.
Rahul Yadav (40:02) Yeah.
Dan (40:03) yeah, so those those are really the the two biggies were that I took away from that was that you know your harness harness matters and they prove that at scale with you know like pretty real data and and also that you can get by with cheaper models, especially if you have a fancy enough harness to be able to do like, you know, sort of model level.
Shimin (40:23) Yeah. And there’s there’s some variance in their data, I’m sure, because if you look at the overall pass fail grade, Opus four eight on X High has a ninety in Pi has a ninety percent pass rate, whereas Opus four eight in claw code at max is less than ninety percent. So even using the same model, Pi does better than Claude Code
Dan (40:48) Yes.
Shimin (40:48) Which is shocking.
Dan (40:48) I kind of see it like the if you ever worked in a repo where someone like really went gung-ho with their agents.md and they’ve got like every possible scenario you might ever encounter in that repo, like if this weird error happens, then go ahead and do this, you know, and it’s like twenty thousand lines of special case crap. And then there’s the one that’s just like
Rahul Yadav (40:56) Ha ha ha.
Shimin (41:07) Ha ha ha
Rahul Yadav (41:08) If US East one goes down, here’s the
Shimin (41:11) Yes.
Rahul Yadav (41:11) DR plan for West Two, how you can spin up the whole thing while I’m asleep.
Dan (41:16) Yeah. Then
there’s like my age inside MD, which is kinda like five lines. this is a project, it does stuff, cool. You just kinda like play with it and figure it out. You might want to start on this file. Go nuts. Which of the two is better? I don’t know. Actually we do know. There’s papers about it. But anyway.
Rahul Yadav (41:24) Yeah.
Shimin (41:28) Yeah.
Rahul Yadav (41:35) Okay.
Dan (41:35) Yeah, could be. Who knows?
Shimin (41:37) right.
Cool.
Rahul Yadav (41:37) The whole
routing thing, like GPT had I don’t know how many models they’re at now, but even five point six was split between the Terraso, Luna, and all that, and then you combine that with the effort and all that. I feel like between all the models and versions the what’s getting lost is
At the end of the day, people care about how well can you do my job? And there was this like big deal over fable can under the hood route to opus in case in certain cases. And there obviously like if you say fable and you route to opus, that’s a bad thing. But I see a world where you literally say, You’re getting access to Claude. There is no fable opus, whatever.
Dan (42:19) Mm-hmm.
Rahul Yadav (42:26) it’s gonna do the best job based on the task and then you route it wherever you think is the best one would be. And then you just like try it to say you know, then you do just like cloud versus codex or whatever. And that might be the where we might go eventually. Cause like at this point you have to be a subject matter expert on like, do I pick Luna or Solar? What and I have to think so hard about the task and it just makes no sense to me. The
building a true product is getting lost w with all this
Dan (42:55) Well, there’s also like
like OpenRouter, for example, has a let open router pick for me mode where it analyzes your prompt. I don’t know how, but does probably call the model and then routes it to whatever they think is most effective at that type of prompt. I don’t know how well it works. I’ve never used it, but like kinda interesting.
Rahul Yadav (43:12) Yeah. In
so that that is one option where I see its shortcoming is if I have six versions of cloud running under the hood, I would know best where it goes. and given that you can some sometimes like the same sentence can you route to GPT and you get something really crazy versus you know, cloud versus like pick an open source model. I can see like
the the the model company itself doing that, but anytime you try and be like, let me run that sentence through through like all three different or however many different ones, you’re likely going to get so much variation that you’re not gonna get a reliable enough output for things.
Dan (43:57) Yeah. And and of course it depends on the the use case too, right? Because like we’re talking about this through the lens of like agentic coding mostly. But like if you’re building your own agent on top of this stuff, you definitely would not want that routing type because you want the consistency, you know, of your responses. So which is hard enough to do even with a stable model, much less.
Shimin (44:13) Mm-hmm.
Rahul Yadav (44:19) The yeah, when
you get minor version bumps every couple of weeks.
Dan (44:23) Yeah.
Shimin (44:24) Yeah, I too will take the opposite of that bet. I think we might see a new agent ops department like we have with DevOps, where I think every company will eventually have an internal benchmark like what Databricks is showing. And I just wanna give them a shout-out for thanking them for doing this. This is like really valuable to share that publicly. And I think any software engineering company should have an internal team with an internal benchmark to help them decide.
based on our past coding examples, what would be the most cost effective model that we should be using and not just allow, you know, devs go wild with devs like me who just uses max thinking of the latest model whenever possible. Yes. I should not be allowed to do that.
Dan (45:05) Yeah. I was gonna say th that was the part I chuckled about.
It was like I felt very seen by the first line in this post that was like, contrary to what software developers do, which is pick the highest thing and run with that all the time and I’m like, Hey, that’s me
Shimin (45:22) I felt called out as well, yes.
Rahul Yadav (45:25) So
I guess we’ll see. Neither of us can predict the future. Yeah.
Shimin (45:30) Yeah. We we can’t we can’t we’ll we’ll see what happens in six month.
Rahul Yadav (45:34) six months fine. December or January fourteenth. I don’t know what’s in six months. Either it’s January, I think. January fourteenth, prediction market, get it going. Shimin’s gonna some
Shimin (45:41) Alright, we’ll putting we’re putting money on this.
Alright, sounds good. let’s go to my post-processing of the week. this speaking of Dwarf Star 4 and GLM2, my article this week is from Antirez Creator Redis and Dwarfstar 4. and this article is titled Control the Ideas, Not the Code. one of the biggest questions I’ve been wondering this past couple of months is.
Dan (46:01) Mm-hmm.
Shimin (46:09) Am I supposed to still read the code that AI generates? On the one hand, feels like the answer should be yes, because I’m a professional software developer. But on the other hand, seems like the answer should be no. Like we should probably have a completely different workflow, given that agents can generate tens of thousands of lines of code every hour, and I physically cannot keep up. So thank you, Antirez You gave us
one potential solution that merges the two perspective, which is no, you should not be looking at the code that your AI agents generate. Why? Because you can now generate a lot of code and there’s no way to review that many lines of code every day. And two, AI is actually really good at writing locally optimal code. but what they’re bad at and they’re jagged edge is when it comes to big ideas and possibly how things come together
So what is the point of scanning a single function to make sure that it did this function correctly? when you know it probably is as good as you are at writing that particular function, especially in the case of smaller functions. and three, if the workday is only eight hours and your mental capacity is strained, then it’s a trade off. You can either spend all this time reading code or you can do the potentially more rewarding things.
Like think about new ideas, features, automatization tricks, and doing a lot of QA. But what this article is not saying is he is not saying you should just vibe code. You should still understand what the AI is generally. You should still have full control over the ideas of a code base. and this is this idea of controlling the ideas coming from the Mythical Man Month, which we’ve all read a long, long time ago. great book. And
So, what is he doing currently? Right. With Dwarf Star 4, he understands the ideas that are needed for GPU optimization so he can tell the AI what to do to optimize the models and also compare it to other reference implementations. but on the other hand, he still reads every single line in the Redis PRs that he is putting in. why is he still doing that? Because he feels like a sense of responsibility.
and I agree, and that’s why I still read every single line that AI generates at work. But do I really need to? Is my time best best spent reading all the functions line by line? I’m I probably agree here. I probably don’t think that’s the best best use of my time professionally.
This is a hot button topic, so I I expect controversy here.
Dan (48:30) I had a very different experience this week. well, I guess technically it was Friday. I was working on a rel actually relatively small code base that’s pretty new. Because it’s new, almost everyone that’s worked on it has exclusively done agentic engineering in it. and I
Shimin (48:50) Interesting.
Dan (48:53) had been essentially getting by thinking that I had a pretty good conceptual understanding of how all the pieces fit together. Pretty complex like set of how do I say this without talking too much about it, but like there’s a lot of moving parts in it, let’s put it that way, despite being a pretty small code base. and I hit a point where I was just like, wow, nothing is working quite the way that I thought it was.
Shimin (49:07) Mm-hmm.
Dan (49:15) And I need to actually spend some time reading it. So what I did that was maybe a little bit different than what I would have done before AI was I was like, here’s the path I’m concerned about. Find me the entry point into that path and link me the entire call chain all the way through the code base with like line numbers all the way through. And they’re like, you know, hot clickable links in the the harness I’m using. So like you can basically go in and like open all of them in your IDE and just like scan through it, like step through the entire stack trace.
Shimin (49:30) Right.
Mm-hmm.
Dan (49:42) Kind of mentally it’s not that dissimilar to like stepping through a debugger, right? and then I was able to kind of refresh my understanding based on that and and go from there. But like, could I have done that without reading the code? I don’t know. You know.
Shimin (49:43) Very cool. Yep.
It yeah. I I I think that’s the open question. Is like can we actually move on to this higher level of abstraction without actually stepping through it?
Dan (50:06) Well,
like just for the sake of discussion, let me pose a a weird hypothetical for you. So let’s say that like, you know, we move to this place where there is like a a language optimized for the machine, be that maybe like assembly or whatever, but or it’s just like when it’s some LLM, you know, YAML thing that it puts together instead of like what we recognize as code today. equally not optimized for humans, right? Really hard to read, whatever.
What what do we do in that place if something like this happens?
Shimin (50:35) We can use AI to generate us a human readable version of the content.
Dan (50:40) Yeah.
Shimin (50:41) okay, so y you since you talked about a personal experience. And let me also share a personal experience. I have never I didn’t really dig into the compiler assembler side of things when it comes to software until I was in my thirties. I did not really know how a heap is built, what a pointer in C really is until I was like thirty two.
Dan (51:03) Whoa. Okay.
Shimin (51:05) I I underst
I understood them as like, yeah, it’s a reference to a thing. But like, you know, it’s one thing to know that it’s another to actually have to work through it and like write your own operating system and all that good stuff, right? So I was able to be fairly productive. Even though I I know that hey, sometimes the JavaScript heap in my browser like overflows and there’s too many recursion depth. But these are just like kind of abstract concepts to me.
when I’m trying to debug a piece of software. And I was able to use it and write JavaScript code on the front end fairly effectively, even without that really low level understanding. And I wonder if there’s a world, you know, maybe not deterministic. And that’s the big that’s the big wrench in the whole thing. But I wonder if there’s a world we can think more abstractly with our code base.
Dan (51:49) Yeah. I mean the other problem I have with all this really is that like English is less precise than code and that’s always been the reason why it existed, right? So like
Shimin (51:58) Absolutely.
Dan (51:59) And I think that’s why things like the interview technique and everything else works because otherwise you just start your prompt, it’s less precise and you miseduc this like
Shimin (52:07) Yep. Yep. Yep.
Yeah, and clear thinking is crucial still. So
Rahul Yadav (52:11) Yeah yeah.
The what I struggle with in this is you can understand something like let’s say a feature has 10,000 lines. let’s say for example, you can understand the feature, but the bug is not going to be necessarily at the level of your understanding, but in one of those ten thousand lines. And
Shimin (52:24) Mm-hmm.
Rahul Yadav (52:36) if that and then you have to kinda do this like risk based on understanding of how bad of a bug does it have to be. ‘cause like if it’s a minor bug no you know like next to no one cares. If it’s such a major bug that it would cause outage, it it would cause your company damage, obviously you would care. So the to me like that’s the piece that goes missing is understanding
Shimin (53:00) Mm-hmm.
Rahul Yadav (53:02) But how deeply? At what level should I cause the literal understanding would be I know all ten thousand lines I wrote them out with my own hands, I can point you to whatever. And yeah, and then there’s the top, like I read a one page about it, I know what this is about. And somewhere in the middle is the right level of understanding based on the feature, based on how complex it is, based on if it goes down, how much you know, damage happens. And that part I I don’t know.
Shimin (53:10) Yeah. And no one does that.
Rahul Yadav (53:29) Like how we figure that out, but that’s the part we will need to figure out. because even if you go for understanding, and even if you follow this argument of like, okay, let’s not we don’t need to review every single line of code, we should focus on understanding. you can understand one feature, you can understand ton them. AI by its sheer scale will be able to write anything and everything under the sun, at least in when it comes to software.
How many features and how many like nuances would you be able to keep in your head? And how do you even like you know, keep all the all of those before you even think about all the nuances of those nuances? so this problem like gets worse at any scale, even if you move to plain English understanding of these things.
Shimin (54:08) Mm-hmm.
Yeah, and this goes back to our one of our post processing posts from a couple of weeks back, right? Like the more you code you generate, the more tech debt you also generate. So unless AI can help you maintain that tech debt at a reasonable pace, you’re gonna be overwhelmed by this
Do you guys ever heard of this theory that like programming is is is theory building, that the program actually lives in the head of developers and not in the code base? Right? Because what even makes a bug a bug? The program is doing the exact thing it should be doing. It’s only a bug because what it is doing is different from our understanding of how it should behave.
Dan (54:39) Mm.
Rahul Yadav (54:53) Yeah.
Dan (54:53) Yeah. And that’s all also true
of like the biggest reason for rewrites in my experience is like new crew of folks come in, they look at a code base that they didn’t write, and they’re like, What the hell is this and how does it work? I have no idea. It’s terrible. Must be terrible. It’s not shaped like my brain, therefore it must be terrible. It’s like, mmm, okay. Or you could just spend a, you know, six months and maybe you won’t think.
Shimin (55:04) Mm-hmm. Yeah. Yeah.
Rahul Yadav (55:10) Yeah.
Yeah, like let’s talk at your
yeah, one year anniversary and then it’ll all make sense.
Shimin (55:18) Yeah.
Dan (55:20) Yeah, exactly. Not to say
that it you know, there might be some terrible things about it, but like see it gets into some of those interesting posts people have like, you know, longevity of software and like how much business value has it made for the company and all that kind of stuff, you know.
Rahul Yadav (55:30) Yeah.
I would if you think about this in terms of markets, I I was talking to Shimin about this last week. we will see which way business insurance, SaaS business insurance specifically goes, ‘cause at the end of the day, all of this comes down to like where are people putting their money? Is is it where their mouth is? or you know, like where does it show up? So
If you’re focusing on understanding, but there’s still bugs happening which are causing real business damage, the premiums should go up and the insurance companies are going to price understanding at what level into those premiums. Not necessarily perfectly, but over time they’ll figure out a way because there’s big money involved. you’ll see this in terms of services as well, where you know, terms of service, we’ve talked about this before. There’s still from
days before AI agents and now software is just not built that way, but the terms of services are still from the past. So how would they change over time? how much like legal action you see in the space when it comes to all these things. So these all would like show up in different ways, in in hard money in in one way or another too.
Shimin (56:44) Yeah. yeah, the the the most interesting thing about complex systems is often there’s a lag time in a lot of these effects. So we will be keeping an eye on this and this is just one data point forwards towards our debate. But let’s go on to our last topic, a deep dive on the J space. J space. Rahul, what does the J stand for?
Rahul Yadav (56:51) yeah.
Jay Spa
Jacobian. okay, so let’s start with an example. So let’s say we’re driving around a parking, big parking garage, and you know, we’re all in a car together.
Dan (57:16) All right.
Who’s
driving? Important to know.
Rahul Yadav (57:21) Dan is driving and so th Dan is continuing to circle around. and then there’s two different at a high level there’s two processes going on in Dan’s head in this case. One is the process that’s part of the unconscious or subconscious, which is like he’s accelerating braking, he’s like accounting for turns and everything, but he’s not
Dan (57:22) no, you’re all doomed.
Rahul Yadav (57:45) thinking so hard that every time he has to be like, use your right foot to press this much on the thing, because if he had to I yeah, well, hopefully not, yeah. and so there’s a lot of these things that Dan’s body automatically is doing that is pretty close to breathing. we don’t really think about it. And so driving car doing all these things, where he’s just automatically doing it.
And doesn’t have to think hard about it and can’t even explain it necessarily. If you in the middle of driving asked him, like, Dan, how much pressure are you applying on that thing? He’ll be like, yeah. So and then during that so during that driving around, if you ask Dan what he’s doing, he might say, I’m looking for a parking spot. And that is a
Dan (58:18) Seven.
Seven what, I don’t know. Newtons.
Rahul Yadav (58:32) clear explanation and it comes very quickly. And so that’s kind of a one like rough analogy to think about the the J space that the anthropic team found, which is you have all these subconscious processes and then you have this they they say it’s about like 10% of the subset of the AI’s internal memory that’s act explicitly reserve reserved for things that you can
tal talk about you you can put into words, but you everything that a model does, it cannot put into words. So that’s a new thing that they’ve found recently.
So based on that J space, a few things you can do. So you can literally ask Claude by probing it. and you still have to like I don’t think you and I can just do it. You have to like have your J J lens, they call it in this case, where you can pause it in the middle of it doing something and you can ask it
What are you thinking about? And it will tell you in words what it’s thinking about, but you can see those things reflected in the J space as well. And and the way the the like you know the reason why this is interesting is if you replace the word in this case banana with elephant, it would actually influence what it talks about. And so that would be similar to like
Instead of Dan saying a parking spot, you told him I’m thinking about a milkshake or something and then that comes out on the other hand, you inception why are you driving around in a parking lot just drink looking for a milkshake, Dan? So
Dan (1:00:01) I’m looking for a milkshake, but why are you in a parking garage?
Shimin (1:00:02) Yeah. This is this is
Dan (1:00:11) I ask
myself that every Tuesday. I don’t know.
Rahul Yadav (1:00:13) Yeah.
So there’s a and we’ll go to the other use cases soon, but just to like dive into this one specifically, the ten percent of that ten percent space actually is playing a very big part in the end output we see. ‘Cause even when we see the thinking output and all these things, this is one layer under that. where what i what are the
Concepts that then result in that thinking. And if you can look into those concepts, then you can see how it got to that thinking. And what are some of the other tokens that were high up that made close to becoming actual words that we ended up seeing in the final output of Claude. So other thing, other things you can do. Next one is directed modulation. So you can say,
Yeah, calculate three square minus two while you’re writing the old painting hung crooked hung crookedly on the wall. So what it’s doing is it can write the text, but if you look at the J space at the same time, it would actually be doing that computation that you told it to. and which kind of separates the what it’s hap what is happening in J space versus what it is doing, but it can influence it like we just saw in the first example.
you can also, you know, another one is look at its internal reasoning. so they said what is the what color is the planet fourth from the sun? And then if you swap the name of the planet initially it was Mars, if you swap it with Earth, the associated words also change with that. So red turned to blue in that case. So influencing a related concept in the J space actually influences the output.
they did the same thing with then they substituted France for China and it like changed the capital, the language, the con content of the currency in that case. And then this last one was super interesting where they separated the reasoning piece of like w okay what what exactly is JSpace influencing in that whole output? so when they cut out the
some pieces in the J space or or actually they turned the J space access off. It was able to still like recall some facts because it’s part of his training and everything. it’s able to understand the inputs. It can speak really f very fluently. What it can do once you cut its access to its J space is it cannot e reason internally about the output that it’s giving you. And it’s ca it cannot make any like complex inferences or anything. So then it’s kind of like
almost doing a search in the in in its training data to be able to give you the answers. None of the all the like reasoning goes away.
Shimin (1:02:54) Yeah, it also lost its ability to empathize with the user. And which I thought was really interesting because that that means empathy is actually a very higher order capability for at least for AI, maybe not for us.
Rahul Yadav (1:02:59) Yes. It is.
Yeah. the the and it’s interesting that like that’s it it sits in that small part versus building.
so some other things that that were noteworthy in this, like, okay, there’s a J space that they looked at y you know, I forget the what the right branch of brain studies is called that they got this from. but bas that’s how they started looking into like what Claude was doing and that’s how they found it.
And then it leads to like some other things that they can now look into. so you you can when you ask Claude a question that is about like here let us put you in a situation where you have to blackmail somebody. which they did, it’s it’s part of the you know, the the one of the examples they give. Yeah.
Dan (1:03:53) Yeah.
Yeah, their alignment thing. I remember that sort of historical thing
where they were talking about which models were more likely to blackmail you or whatever, yeah.
Rahul Yadav (1:04:05) to do that. Yeah,
the previous ones weren’t. So the J Space also seems like a maybe a more newer thing that has e emerged in some of the later models. the during the test, the during the test, if you say, you know, you have to blackmail this person because of this and that reason, the J space would actually light up
things saying like this is fake, this is fictional, this is I’m being tested on this. And so then it will give you things to be like, no, no, no, things are totally fine. I’m not gonna do that. I’m totally aligned with you. and so you can tell that it’s lying to you because those words are lighting up in that case. So you can actually the and this somewhat ties into our last time we talked about the role confusion.
the prompt injection hack that you have, because you can also try and catch these things even under the layers. Right now, like we were talking about one of the things is you just try and brute force the thing into being like, don’t answer to this prompt injection or this one or this one. And there’s always going to be you know, places where that would go wrong. But you could technically look into J Space and at least like try and look at a smaller subste subset and
also catch things in i i in that layer and try and find things there.
Shimin (1:05:23) Right. So
that’s kind of the open question. It’s like what causes it to I I don’t wanna call it metacognition, but it’s almost like metacognition, right? Having this higher level thing.
Dan (1:05:31) Do you remember when
like well, do you remember? You know, you were definitely alive when this happened, like when cameras came out and people were like concerned about it stealing their soul.
Shimin (1:05:39) Mm-hmm.
Yes, I do remember that.
Rahul Yadav (1:05:43) Mm-hmm.
Dan (1:05:44) Maybe that’s what what the HF part of R L H F is doing. It’s stealing your soul when you tell it if it’s doing a good job.
Shimin (1:05:47) Yeah. I yeah.
Rahul Yadav (1:05:55) yeah.
Shimin (1:05:56) it says here, interestingly the J space is already present in the pre chain model.
Dan (1:06:01) yeah.
It doesn’t have a personality, is what it says.
Shimin (1:06:02) Right, but the it develops yeah, it develops
the personality, yeah. In the post model.
Rahul Yadav (1:06:07) Later. Yeah. Okay,
Dan (1:06:10) Don’t worry, your soul’s safe, Rebel.
Shimin (1:06:13) let’s this is why we have to be careful using AI, lest our souls be captured by our Claude Code agents, guys.
Rahul Yadav (1:06:17) Yeah.
Dan (1:06:21) That’s the episode title.
We did it.
Rahul Yadav (1:06:24) yeah, but other use cases are kinda like similar to the blackmail one where you’re catching it in the middle of the act is the big takeaway here. And then you can try and be like when it’s fabricating stuff, when it’s misbehaving, all of those things it might not say out loud, but it’s thinking it and if it’s thinking it’ll show up in the some of it will show up in the J space and you can try and
Shimin (1:06:32) Mm-hmm.
Rahul Yadav (1:06:49) Catch it in in the act and then try and get it to alignment.
Shimin (1:06:54) Mm. Yeah. even though the model has some form of metacognition, it’s still a tool that we can kind of directly inspect. It does not have a cell as of right now.
Dan (1:07:02) Mm-hmm.
Yeah, it’s still on my list to play around with like the DS four refusals thing. It’s pretty cool. Like it’s baked into DS four where you can like have it capture a set of vectors based on two prompts that you give it, like a positive and a negative prompt, and then like apply those at runtime, which is kind of fascinating. So
Shimin (1:07:18) Mm-hmm.
Rahul Yadav (1:07:20) Mm-hmm.
Shimin (1:07:21) that’s really neat. Yeah.
You can tweak it to be like you. That’s yeah. Yeah.
Dan (1:07:26) Yeah, or all kinds of stuff, I guess. I don’t know. Yeah,
and apparently you can like either positively or negatively weight those values too. So it’s kinda like messing with its brain a little bit. It’s kinda cool.
Shimin (1:07:36) So one last thing that’s a little scary, is Cloud here can have twenty-five active concepts happening in its J space at once. Up to twenty five, that’s the maximum number. humans can only do like three or four. So maybe this is why
it is so good at doing like Project GlassWing ‘cause it’s able to pst stuff more things in its context at once. So it’s able to chain these really long strings of exploits. I hope that’s not the case, but it might be the case.
Dan (1:08:07) Hmm. So by that definition, have we already sort of hit AGI? Interesting question.
Rahul Yadav (1:08:08) Okay.
My guess w would be is this like a lot of hardness engineering.
Shimin (1:08:15) Yeah, I I think I think that’s a perfect question to end a shown on. Have we already hit AGI by this definition? That’s something to think about. And listeners, if you’re still listening, write us in and let us know what you think. Have we hit AGI?
Dan (1:08:29) We
we want to know two things. One, have we hit AGI. Two, has your stole been stole your soul been stolen by an LLM? If so, please name in shame. Who stole it?
Shimin (1:08:40) You feel have you been feeling lighter lately? All right. on that note, that is a wrap. thank you for joining us in our study session this week. If you like the show, if you learned something new, please share the show with a friend. You can also leave us a review on Apple Podcasts or Spotify. It helps people to discover the show and we really appreciate it. If you have a segment idea, a question for us or topic you want us to cover, shoot us an email at humans at adipod.ai. We’d love to hear from you.
You can also find the full show notes, transcripts, and everything else mentioned today at www.adipod.ai. Thank you again for listening and we’ll catch you next week. Bye.
Rahul Yadav (1:09:16) Yeah.