Hi! This is a new benchmark that I created together with other researchers at Meta, Stanford, Harvard, UW. Most benchmarks these days seem to test models to just fix a bug that I as a user already encountered. But shouldn't we expect models by now to also find bugs before anyone runs into them? So…
We're a team of 4 and we built Lemma, an open-source workspace where people and AI agents work in the same place. I work on it (I do GTM, not engineering). **How Claude was used to build it** Most of Lemma's code was written by Claude. The process is the same for every change: * A human writes the…
95% of the functionality was one shot with Opus 5.5 on high. Described what I wanted and it just worked. https://imgur.com/a/PCiM6Wj https://imgur.com/a/PGb8jsH (Just a demonstration video from Claude, the extension pulls icons from IMDB) https://github.com/delphizer/reel-weights IMDb's search…
If you've used Gonka inference directly, you've probably noticed individual brokers aren't always consistent — one might be fast and reliable for a while, then slow down or drop requests, then recover. That's just how a decentralized network of independent nodes behaves. SAGG takes a different…
This started as a tiny Python script to screen NSE stocks for myself. Then I wanted charts. Then fundamentals. Then I got tired of reading 300-page annual reports... and somehow it turned into a full terminal. So here we are. The video is the part I'm most excited about: I ask the built-in agent…
a small update to this digital wellbeing software since my last post. This is a prerelease version which i have tagged to be stable and slated for future release and add to main branch from the beta branch. This highlight is less about asking you guys to try out the beta release but more in the…