hlfshell Keith Chester

ai

I am unreasonably excited about Taalas

I find myself unreasonably excited for the work that Taalas is doing. They’re turning open weight models directly into ASICs (application specific integrated circuits) so that the chip essentially acts as the model - and only as the model. You get raw transistor switching speed within the model. Not only does this result in incredibly fast computation (~15k tokens a second for llama 8b!) but you are doing so at ridiculously low power consumption and efficiency.

I’ve heard some rumors that are tempting and I hope are true. First - a Qwen 27B model is in the works; this is a great model that can do a lot of the simpler tasks you’d want from local AI, and can even do a bit of coding well. Importantly it’s a very strong tool calling agent, so it could very easily hand off tasks to additional AIs when needing capable compute. Supposedly cost to produce the board is ~$400 (again - unstantiated rumors here) which would mean an attractive $800 price tag.

If we start to look at large models - 70b and 120b open models - being translated to this style of chip, we could be looking at an explosion of local AI that completely changes the costs (both monetary and power consumption) of AI application. But…

I truly don’t expect that they will ever make it to selling consumer boards; I fully expect them to either find themselves gobbled up by NVIDIA to secure potential competition, or for someone like Google/Samsung/Apple to buy them to build out the equivalent for phones.

Firecracker + poprocks (DEVx)

I just gave a talk at DEVx giving a very early alpha preview of poprocks, a module providing simple golang primitives for building and working with Firecracker microvms. While it can be used serverside, I talk about utilizing it instead as a client side solution for isolated, malicious actor “safe” execution environments - especially for AI agents.

What comes after the token discount bubble pops?

Coding agents are amazing - sure. I’m a fan and heavy user of them too, especially CLI agents. BUT - the existing coding subscriptions heavily discount token usage, to the point that most users are unaware of just how many tokens they are actually burning at any given moment to run their agents. This blissful ignorance is perfectly fine as long as the discounsts continue; the pain of running the agents aren’t being felt by the end user.

Why are they so heavily discounted? There’s a race to secure funding and users by the big AI model providers, resulting in heavily (prohibitively and unsustainably) discounted tokens to try and build a user base with dependence on them as a provider. Either users will have no choice but to deal with price hikes, or one of these providers will be the “last man standing” and finally make good on their untenable investor promises.

When the AI economic bubble pops for these providers (which, barring any major discovery in terms of hardware / model architecture, will happen), we’re likely going to see a large increase in token pricing. The next step? I see three possible outcomes.

1 - Tooling favors mixed model approaches, where specially tuned agents appropriately route to different model sizes of varying costs based on user preferences for costs and task categorization. We don’t need Claude Opus to center a div, but we do need it for a very complex architectural change. It wouldn’t surprise me that, for internal cost saving, this approach is implemented by the providers themselves as a singular packaged model endpoint or incorporated into the tooling.

2 - We see a sharp rise of business-focused large scale token pricing packages, similar to how we see reserved instance pricing in cloud infrastructure. Buying a years worth of tokens up front based on projected dev usage nets businesses discounts. This doesn’t bode great for model providers though - it creates a larger scale race to the bottom.

3 - Developers and product builders start buying the best possible pro-sumer hardware for headless agent machines (current day equivalent at time of posting would be something like an AMD Strix Halo Processor, like the Framework Desktop (which can run heavily quantized 70B models) ) to run smaller models that can do MOST work, but can call out or pass over to a more expensive model as need be. A kind of local model play on # 1.

I’m personally hoping for some mixture of 3 and 2. I would love to be able to run models locally, but understand that the hardware scale is a long way out from becoming cheaper or consumer grade. Being able to reserve token pricing in bulk to hand off on “capable enough” models to occasionally “upscale” is likely key.

Plus I just prefer local, user owned compute.

ARC-AGI-3 beta is live!

ARC-AGI-3 beta is live!

For the past few months I’ve been working for the ARC Prize - a non profit organization built around the idea of an Abstract and Reasoning Corpus - a set of reasoning games that are easy for humans to quickly and efficiently figure out, play, and solve - but near impossible for even the most cutting edge models. [Paper]

…and today we officially launched the beta of our toolkit and three games for the public to try out!

For ARC-AGI-3, we’ve prepped over 150 games to test AI. I’ve been building out tooling (most notably the benchmarking agent tools) to make it dead simple to test and research models and various agentic architectures.

Using this tooling (and I have more to release soon!) I’ve also been researching how models reason about these games. I’m trying to develop new architectures and techniques to maximize performance of the models and zero in on how these models reason internally and understand uniquely abstract state spaces. This is acutally quite difficult; these games are deceptively simple - outright easy for humans - but even so-called-superhuman LLMs and AI products can’t solve a single one of them.

Check it out and feel free to reach out to me to chat about it.

threadsafe_datastore

Just released threadsafe_datastore, a simple, convenient thread-safe data store for Python.

I kept rebuilding this feature to pass around context within AI agents working on the same data across multiple threads. I kept wanting a simple atomic datastore that was convenient to work with. I originally built it for arkaine | git |. So - here it is as a stand alone package for easier use. I also improved context management for real easy multi-step operations in case you need to do something real custom.

from threadsafe_datastore import Datastore

store = Datastore()
store["counter"] = 0
store.increment("counter", 5)  # Returns 5

# Nested dictionary support
store["nested"] = {"items": []}
store.append(["nested", "items"], "value1")

# Thread safe multi-step w/ context management:
with store as unlocked:
    unlocked["a"] = 1

    # This would be unsafe without the context:
    unlocked["b"] = unlocked.get("a", 0) + 1

Give it a try: pip install threadsafe-datastore

structured-parse release

Just released structured-parse, a multi-language parser for block labeled LLM output. It’s based on the parser I had built for arkaine.

Here’s what I mean by “block labeleled” output:

Thought: I need to search for information about robots
Action: search
Action Input: {"query": "robots and why they're so cool", "max_results": 5}

This output is not only more human readable, but also easier for LLMs to produce. But there’s a catch - LLMs tend to still introduce nondeterministic volatility towards these outputs; humans are just good about reading through that. structured-parse is a robust parser that can deal with this, allowing LLMs to reliably follow instructions and allowing your code to parse it into clean, typed data structures.

structured-parse is written in Go with exports to TypeScript/JavaScript and Python via WebAssembly; so it’s all golang at its core.

Give it a try!

Missing arkaine already

Recently I took on a contract job to fill the coffers a bit while HiredCoach handles its initial launch. Nothing that I can talk that much about publicly, but I can reveal that it’s your typical AI office assistant for a particular business process. As specific as it is exciting a description, I’m sure.

Given that I recently wrote an entire article about the highs and lows of my custom framework, arkaine - and within I mention that I lkely won’t continue developemnt on that particular implementation of those ideas I wouldn’t exactly want to pull it out for a client hoping for their PoC to be easily maintainable for future developers as they pursue it; it’s too esoteric a framework and oto tied to my own development preferences.

That being said, I still have difficulty finding a replacement that works in the way I enjoy or rate highly. I find myself wishing for several of the features and niceties I baked into arkaine. I view it as a reassuring sign that I was onto something with a few ideas there.

As an additional note - Web research agents that are useful are hard to get right. I’ve built these three times now and still feel like they are tricky to make them reliable and performant.

I started working on “arkaine 2.0” (which will likely be the 0.1.0); whether I continue on it remains to be seen, as it is architectually a gigantic change. I’m playing with Merkle trees and more controlled syncing paired with better serialization and compression, as well sa more flexible composable functions without falling into the overly verbose repetitive context nesting we ran into before.

The Physical Turing Test: Nvidia's Vision for Embodied AI

It’s certaintly a fascinating time for robotics. I am generally pessimistic of the current state, and likely fates, of much of the current landscape of the robotics industry, but it is undeniable we are rapidly unlocking new capabilities and research is flying forward.

This talk acts as an excellent high level overview of the key innovations to our approach of utilizing reinforcement learning for better performance of robots with our VLMs / VLAs for a more general audience.

If you want a slightly deeper dive (though admittingly out of date with the pace the field has moved) checkout my overview of LLM research aligned with robotics from 2023, and my thesis work of using LLMs to utilize heuristic knowledge to understand context of requests and environments. If I were to do a similar project today, I would certainly take a more agentic route. I certainly wish I had written arkaine prior to that project - it would have avoided so many headaches.

go-arkaine-parser

Back when I was working on coppermind at the heyday of GPT3.5’s initial world shattering release, I had… difficulty finding good ways to deal with parsing the stochastic LLM outputs.

I got better at this when I started work on arkaine, eventually developing a pretty useful and reliable parsing pattern.

With an idea that would be best served as a golang app requiring interacting with LLMs I decided to do a quick port of the parser to an idiomatic golang module.

So if you need AI parsing for your golang project, check out go-arkaine-parser.