<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="/feed.xml" rel="self" type="application/atom+xml" /><link href="/" rel="alternate" type="text/html" /><updated>2026-07-30T05:17:59+00:00</updated><id>/feed.xml</id><title type="html">Andy Konwinski</title><subtitle>Notes on AI, RL, systems, etc.</subtitle><author><name>Andy Konwinski</name></author><entry><title type="html">Concentration of power in AI is a risk, not a solution</title><link href="/2026/07/02/concentration-of-power-in-ai-is-a-risk-not-a-solution.html" rel="alternate" type="text/html" title="Concentration of power in AI is a risk, not a solution" /><published>2026-07-02T00:00:00+00:00</published><updated>2026-07-02T00:00:00+00:00</updated><id>/2026/07/02/concentration-of-power-in-ai-is-a-risk-not-a-solution</id><content type="html" xml:base="/2026/07/02/concentration-of-power-in-ai-is-a-risk-not-a-solution.html"><![CDATA[<p><em>The answer to concentration of power in AI is not openness at all costs, but a serious research commons with the resources to compete.</em></p>

<p>Over the past few months, I’ve become increasingly uneasy about the concentration of power in the AI ecosystem.</p>

<p>Not because the risks of advanced AI are imaginary. They are not. More capable systems could enable cyberattacks, biological weapons development, and other civilizational dangers. If you take those risks seriously, as I do, the instinct to centralize access can seem reasonable.</p>

<p>But that instinct has become one of the least examined assumptions in the AI debate. In the name of safety, we are drifting toward a world in which a handful of private companies decide who can work at the frontier, which research paths are permissible, and which parts of the technology remain visible to the rest of society.</p>

<p>That is a mistake. We need to treat concentration of power in AI as a risk, not a solution.</p>

<p>The dominant conversation in AI safety has focused on misuse: how to stop bad actors from gaining access to powerful systems. That is critical. But it is not the only question. Just as important is how the power of this technology gets distributed.</p>

<p>Artificial intelligence is becoming a foundational productive resource. The right analogies are railroads, telecommunications, oil, electricity, the internet. These technologies did more than expand economic possibility. They reordered society around new infrastructure, shifting who could participate, innovate, benefit and accumulate power.</p>

<p>AI will do the same. Who gets access? Who gets to build? Who gets to participate? Who gets to decide? That is what is actually being debated, whether we acknowledge it or not.</p>

<p>Earlier this month, Anthropic came under criticism for a policy that could silently degrade some of their AI product’s responses when users appeared to be using the model to train competing AI systems. Many in the research community criticized the approach for altering model behavior without informing users. Anthropic reversed course, but the episode revealed something deeper.</p>

<p>When an institution controls access to the frontier of innovation, it can begin to see itself not merely as a builder of a foundational technology, but as governor of it. The problem isn’t that they made a bad decision. The problem is that they assumed the decision was theirs to make.</p>

<p>To be clear, many of the people leading today’s frontier labs are thoughtful, serious people trying to navigate genuinely difficult problems. But good intentions are not a governance system.</p>

<p><strong>Democracy is built on a profound skepticism of concentrated power. Open science shares this principle. Both are built on the idea that progress and legitimacy emerge from broad, distributed participation rather than concentrated, gated authority.</strong></p>

<p>What worries me most is that we’re drifting toward concentration not by design, but because of the incentives of a few organizations and individuals.</p>

<p>The resources required to work at the frontier of AI are unlike anything modern industry has seen before. The all-in cost to train a new version of a frontier model reaches into the billions. Access to frontier-scale compute resources is increasingly scarce. The universities that produced many of the foundational breakthroughs behind modern AI are essentially unable to access state-of-the-art systems built on that very research.</p>

<p>Meanwhile, the frontier itself is becoming concentrated behind the walls of a few private companies. The most capable models, the largest compute clusters, and many of the field’s most talented researchers are increasingly gathered inside a small number of closed labs.</p>

<p>We’ve seen a version of this story before.</p>

<p>The internet began as a research network built on open protocols. But that openness was not inevitable. The early analogues to today’s Anthropic, OpenAI, and Google were not consumer platforms. There were telecom and computing incumbents with every incentive to make the infrastructure of the internet a closed system controlled by a few dominant actors.</p>

<p>It didn’t happen. Not because everyone agreed openness was morally superior, but because specific people made specific decisions at critical moments to keep foundational infrastructure permissionless and interoperable. Vint Cerf helped develop the protocols that allowed networks to communicate with one another; Tim Berners-Lee later designed the web as a system anyone could build on. From there, researchers built on shared foundations. Entrepreneurs created products without fear of arbitrary exclusion or negotiating for permission. Entire industries emerged on top of infrastructure that no single company owned.</p>

<p>The result was not simply a better internet. It was an explosion of innovation no single company could have planned. Cloud computing, the modern software industry, much of the data economy, and even the training data behind today’s AI systems all stand on top of the decision to keep the underlying infrastructure of the internet broadly accessible.</p>

<p>This is the lesson for AI. Today, we are making different decisions, many rational, most well-intentioned, and all compounding.</p>

<p><strong>If our best scientists and engineers can only reach the frontier by joining a handful of secretive labs, we do not have an open research ecosystem. We do not have a truly competitive market. We have a system in which participation increasingly depends on the permission of a few individuals at a small number of private companies.</strong></p>

<p>We are moving rapidly toward a permissioned frontier.</p>

<p>And that matters for innovation. It matters for economic agility. It matters for national competitiveness. And yes, it matters for safety.</p>

<p>If advanced AI could create species-scale risks, then more of our best researchers should be working to understand them. Safety is not secrecy. Safety is error correction. Modern medicine, aviation, and engineering progressed because talented people could openly challenge assumptions, discover mistakes, and improve the state of the art. Broad participation is how we avoid catastrophic failure. It is also how we make progress.</p>

<p>We do not have to accept a future in which the AI frontier is accessible only inside a handful of private companies. But warning about concentration is not enough; we must build a credible alternative.</p>

<p>That means creating a heavyweight contender in the open: a research commons that brings together the world’s best minds across institutions and gives them access to the resources to work at the frontier. Much of the innovation that powers today’s AI emerged from universities and open scientific collaboration. But the frontier has outgrown any single university. The challenge now is to build a new commons at the intersection of academia, industry, and the public interest.</p>

<p>This research commons must be ambitious enough to matter. It will require frontier-scale compute, access to state-of-the-art models, top talent, operational support, public investment, philanthropic capital, deep collaboration, and companies willing to contribute to an ecosystem larger than themselves.</p>

<p>Our top researchers should not have to choose between independence and relevance; they must be able to reach the frontier without joining a handful of private companies. Academic contributions should not exist solely to improve the products of private firms. And the path from scientific discovery to entrepreneurship should remain open.</p>

<p>The goal is not to recreate the frontier labs. Even a well-resourced open commons will not match the budgets of the largest private AI labs. That relative scarcity can spur breakthrough architectures, alignment techniques, evaluations, and other research paths that frontier labs may overlook.</p>

<p>Yesterday, I spent a long day at Open Frontier, a meeting with many of the top minds working on open frontier AI in academic and industry labs, scientists and engineers who will not be spectators while the frontier of human knowledge is pulled behind closed doors. At the gathering, we resolved to manifest the open research commons outlined above and keep the frontier open (much of the day’s content was livestreamed <a href="https://www.youtube.com/live/-41kYH6JgvU">here</a>).</p>

<p>If we get this right, the benefits of intelligence will not be confined to the few entities that increasingly control it. They will diffuse outward, into new discoveries, new companies, new industries, and new possibilities we cannot yet imagine.</p>

<p>The frontier will advance. What remains undecided is who gets to advance with it.</p>

<p><em>I also posted this essay <a href="https://x.com/andykonwinski/status/2072830533739192560">on Twitter</a></em></p>]]></content><author><name>Andy Konwinski</name></author><summary type="html"><![CDATA[The answer to concentration of power in AI is not openness at all costs, but a serious research commons with the resources to compete.]]></summary></entry><entry><title type="html">K Prize round one closed; what’s next</title><link href="/2025/03/12/kprize-next-steps.html" rel="alternate" type="text/html" title="K Prize round one closed; what’s next" /><published>2025-03-12T00:00:00+00:00</published><updated>2025-03-12T00:00:00+00:00</updated><id>/2025/03/12/kprize-next-steps</id><content type="html" xml:base="/2025/03/12/kprize-next-steps.html"><![CDATA[<p>K Prize Round 1 is Officially Closed - thank you! <br />
<img style="padding:2px; float:right" src="/assets/img/kprize-dollars.gif" alt="" width="45%" /></p>

<p>Well, that was fun. Submissions for the first phase of Konwinski Prize are now closed - thanks to the 616 teams who entered what I’ll call “round 1” of K Prize. The competition was fierce, and this is not an easy benchmark. Congrats to those of you who made it to the Kaggle leaderboard.</p>

<p>Now that the submission deadline has passed, we’ll begin collecting our test set to figure out real scores. We’ll announce a first batch of prize winners on June 11th, so stay tuned. No submissions will be allowed during the competition downtime.</p>

<p>Rest assured, I still intend to give a million bucks to the first open source AI that breaks 90% on this benchmark.</p>

<p>K Prize was meant to do three things:</p>

<ul>
  <li>Measure how AI coders perform when they can’t cheat</li>
  <li>Model a better way to benchmark (contamination free with a prize!)</li>
  <li>Move AI forward in the open</li>
</ul>

<p>And we’re not done yet! We’re already working on the next phase of K Prize, incorporating learnings from this time around.</p>

<p>See also:<br />
<a href="https://kprize.ai/">K Prize website</a><br />
<a href="https://www.kaggle.com/competitions/konwinski-prize">K Prize Kaggle competition</a><br />
<a href="https://docs.google.com/presentation/d/1yp8xWBLB9Uf6VwDFk6SznreUUsegK78jHBAVCuVPAVw/edit?usp=sharing">K Prize NeurIPS announcement slides</a><br />
<a href="https://x.com/andykonwinski/status/1867015050403385674">My tweet from on stage at NeurIPS</a></p>]]></content><author><name>Andy Konwinski</name></author><summary type="html"><![CDATA[K Prize round one submissions are closed]]></summary></entry><entry><title type="html">The first Laude Salon: on researchers shaping AI’s impact</title><link href="/2025/02/24/laude-salon.html" rel="alternate" type="text/html" title="The first Laude Salon: on researchers shaping AI’s impact" /><published>2025-02-24T00:00:00+00:00</published><updated>2025-02-24T00:00:00+00:00</updated><id>/2025/02/24/laude-salon</id><content type="html" xml:base="/2025/02/24/laude-salon.html"><![CDATA[<p><img style="padding:2px" src="/assets/img/salon-1.jpg" alt="A black and white photo of a Dave Patterson on stage kicking off the Salon" width="45%" />
<img style="padding:2px" src="/assets/img/salon-2.jpg" alt="A warmly lit photo of attendees at tables during a salon event" width="45%" />
<img style="padding:2px" src="/assets/img/salon-7.jpg" alt="Me chiming in" width="45%" />
<img style="padding:2px" src="/assets/img/salon-5.jpg" alt="Jeff Dean engaged in the salon discussion" width="45%" /></p>

<p>Earlier this month, I helped assemble some of the smartest folks on the planet to discuss how researchers might steer the AI agenda to a positive impact on billions of lives. The event was the first of many Laude Salons and it was a bit of an experiment - an attempt to foster dialogue among researchers about where AI is heading and how we can shape it for the better.</p>

<p>The evening was co-hosted by <a href="https://laude.vc/">Laude</a> and my co-authors of <a href="https://shapingai.com/"><em>Shaping AI’s Impact on Billions of Lives</em></a> (<a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Cu%C3%A9llar,+M">Mariano-Florentino Cuéllar</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Dean,+J">Jeff Dean</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Doshi-Velez,+F">Finale Doshi-Velez</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Hennessy,+J">John Hennessy</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Koyejo,+S">Sanmi Koyejo</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Moiloa,+P">Pelonomi Moiloa</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Pierson,+E">Emma Pierson</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Patterson,+D">David Patterson</a>). The paper itself had been conceived almost exactly a year before over pints of Guinness with Dave Patterson (one of my heroes) at a Berkeley pub. That day, we sketched out a simple idea: computer scientists need to take a more active role, not just in building AI, but in steering its impact. Rather than merely predicting what AI might do under a laissez-faire approach, we asked: <em>What could AI do if we directed our efforts toward maximizing the upsides and minimizing the downsides?</em></p>

<p>From that discussion, a plan took shape. Over the following months, our team grew to nine of the world’s leading computer scientists and rising AI stars from academia, startups, and big tech. Together, we set out to explore AI’s pragmatic near-term impact—not just in theory, but through real conversations with those on the frontlines of change.</p>

<p>We spoke with two dozen experts across different domains, including:</p>

<ul>
  <li><strong>John Jumper</strong>, a Nobel Prize winner in chemistry, on AI’s potential in scientific discovery.</li>
  <li><strong>President Barack Obama</strong>, on governance and the role of policymakers in shaping AI’s future.</li>
  <li><strong>Susan Rice</strong>, former UN ambassador and national security adviser, on the implications for security.</li>
  <li><strong>Eric Schmidt</strong>, former Google CEO and philanthropist, on AI in the economic and technological landscape.</li>
  <li><strong>Neal Stephenson</strong>, renowned author and futurist, on the impact of AI systems on shaping entertainment.</li>
</ul>

<p>During the Salon, my co-authors and I took turns moderating sections of the conversation (special thanks to <a href="https://obamawhitehouse.archives.gov/blog/author/jason-goldman">Jason Goldman</a> for subbing in to lead our policy segment). The attendees debated topics from the paper, including how to select research moonshots and how to maximize research impact, e.g. through startups and open source. We also had an impassioned discussion about how (or if) we might move beyond H-index as a metric of academic impact.</p>

<p>The evening was inspiring, a bit sobering, and I want more like it. I’ve been thinking a lot about the importance of open discourse lately, and I plan to make sure more conversations like this one are happening.</p>]]></content><author><name>Andy Konwinski</name></author><summary type="html"><![CDATA[A conversation about how researchers can shape the AI agenda]]></summary></entry><entry><title type="html">Why I built the Konwinski Prize</title><link href="/2024/12/12/konwinski-prize.html" rel="alternate" type="text/html" title="Why I built the Konwinski Prize" /><published>2024-12-12T00:00:00+00:00</published><updated>2024-12-12T00:00:00+00:00</updated><id>/2024/12/12/konwinski-prize</id><content type="html" xml:base="/2024/12/12/konwinski-prize.html"><![CDATA[<center style="padding-bottom:20px">
  <div style="overflow: hidden;">
    <div style="float: left; width: 50%; text-align: center; min-width: 350px;">
      <a href="https://x.com/andykonwinski/status/1867015050403385674">
        <img alt="Tweet: I'll give $1M to the first open source AI that gets 90% on this sweet new contamination-free version of SWE-bench - http://kprize.ai" src="/assets/img/kprize-tweet.png" style="width: 90%; padding-bottom: 30px;" />
      </a>
    </div>
    <div style="float: left; width: 50%; text-align: center; min-width: 350px;">
      <img src="/assets/img/kprize-launch-on-stage.jpg" style="width: 90%;" />
      <div style="margin: auto; text-align: center; padding-top: 10px; font-size:.9em; font-style: italic; width: 90%;">
        Live tweeting the announcement at NeurIPS. Left to right: Kaggle CEO <a href="https://scholar.google.com/citations?hl=en&amp;user=l_O64B8AAAAJ&amp;view_op=list_works&amp;sortby=pubdate">D. Scully</a>, SWE-bench creators <a href="https://www.carlosejimenez.com/">Carlos Jimenez</a> &amp; <a href="https://john-b-yang.github.io/">John Yang</a>; and me
      </div>
    </div>
  </div>
</center>

<p>Last night, I <a href="https://x.com/andykonwinski/status/1867015050403385674">tweeted</a> on stage at NeurIPS that I’ll give $1M to the open source AI that breaks 90% on a sweet new version of my favorite benchmark, <a href="http://swebench.com">SWE-bench</a>. The competition is called the <a href="http://kprize.ai">K Prize</a>!</p>

<p>The K Prize aims to:</p>

<ul>
  <li>Measure how AI coders perform when they can’t cheat</li>
  <li>Model a better way to benchmark</li>
  <li>Move AI forward in the open</li>
</ul>

<p><strong>Why SWE-bench?</strong> Well, I fell in love with SWE-bench the moment I saw it. What a great idea: have AI solve real issues from popular GitHub repos. I love that SWE-bench is hard, built from real world data, and measures something that I care about (coding).</p>

<p><strong>Why make a contamination free version?</strong> SWE-bench publishes its test set. This allows leaderboard submissions to train on the test data. Even if SWE-bench didn’t officially publish the test set, the benchmark is composed of issues and code scraped from public GitHub repos, and most models today are trained extensively on those same repositories, so contamination is a likely happening. I’ve always wondered how the leaderboard would change if the test set weren’t public. That’s why I decided to build a contamination-free version of SWE-bench.</p>

<p><strong>Why open source code and open weight models only?</strong> I am passionate about open source. One of the things I love most about open source is how it taps into a bigger community energy - individuals building off the work of other individuals. I love that feeling of getting swept up into something much bigger than yourself. That moment when you look at your team and say, “whoa, are we actually doing this?” That’s how I felt when I was working on Apache Spark at Berkeley, and that’s how I hope those who work on the K Prize will feel.</p>

<p><strong>Why a competition?</strong> I came to appreciate the power of programming competitions to catalyze research progress back at UC Berkeley, where I witnessed the Netflix Prize (which was also $1M and based on real data) inspire my friend and Databricks co-founder Matei Zaharia to create Apache Spark. Not to mention, I love competitive programming myself. And in my experience picking co-founders, competitive programmers make some of the best - Matei and my Perplexity co-founder Johnny Ho were both some of the best ranked in the world.</p>

<p>One day in early 2024 I had the idea to put all of the above together and the K Prize was born! I assembled a team to adapt SWE-bench’s collection pipeline and implement the evaluation part on Kaggle’s infrastructure. For this competition we will collect a new test set after the submission deadline.</p>

<p>And finally, shoutouts to the core members of the K Prize team for standing this up: Chris Rytting, Justin Fiedler, Alex Shaw, K. Tighe, and Lindsey Gregory. Thanks to our creative team for making it cool: Jennifer DeVoid on dev, Richard McClellan on design, and Travis Nichols on Dr. K (aka cartoon me). Also thanks to Kaggle for their deep partnership. Finally, thanks to John Yang and Carlos E. Jimenez (the creators of SWE-bench) for their advice and help as we built the collection pipeline for contamination-free SWE-bench.</p>

<p>Happy hacking!</p>

<p>–</p>

<p>See also:<br />
<a href="https://kprize.ai">K Prize website</a><br />
<a href="https://www.kaggle.com/competitions/konwinski-prize">K Prize Kaggle competition</a><br />
<a href="https://docs.google.com/presentation/d/1yp8xWBLB9Uf6VwDFk6SznreUUsegK78jHBAVCuVPAVw/edit?usp=sharing">K Prize NeurIPS announcement slides</a><br />
<a href="https://x.com/andykonwinski/status/1867015050403385674">My tweet from on stage at NeurIPS</a></p>]]></content><author><name>Andy Konwinski</name></author><summary type="html"><![CDATA[Why I launched the K Prize, aka, why want to give $1M to the first open source AI that gets 90% on contamination-free SWE-bench.]]></summary></entry><entry><title type="html">Laude Ventures</title><link href="/2024/12/10/introducing-laude-ventures.html" rel="alternate" type="text/html" title="Laude Ventures" /><published>2024-12-10T00:00:00+00:00</published><updated>2024-12-10T00:00:00+00:00</updated><id>/2024/12/10/introducing-laude-ventures</id><content type="html" xml:base="/2024/12/10/introducing-laude-ventures.html"><![CDATA[<p>A few years ago, I co-founded a small fund called CSGV (CS Grad Ventures) that took money from — and invested it in — Computer Science PhDs and Professors. Our vision was to democratize Venture Capital back to the researchers themselves. To succeed we would need to find the next Databricks.</p>

<p>That fund is how I met Aravind Srinivas, whom I invested in and co-founded Perplexity with. It has been inspiring to see how the network of research founders has helped Perplexity and how Perplexity has benefitted in return.</p>

<p>Aravind and I were introduced by one of the CSGV PhD LP investors (an AI researcher at Meta). Aravind and I had a bunch of mutual connections that we used to reference each other (e.g., I was co-teaching a class at Berkeley about startups with his PhD advisor). And another of the CSGV LPs (a Berkeley PhD himself) became Perplexity’s founding engineer.</p>

<p>Since my co-founder Andrew Krioukov and I started CSGV, we’ve continued to meet, teach, and advise dozens of PhDs and Professors out of top CS programs that are obsessed with impact. I see my Databricks and Perplexity founders in many of them.</p>

<p>Today, I’m excited to <a href="https://laude.vc/news/from-breakthrough-research-to-breakout-companies/">announce Laude Ventures Fund I</a>, the next chapter for CSGV. The name Laude rhymes with “awed”, and is inspired by Cum Laude (to graduate “with honors”). Not only have our researcher LPs joined us from CSGV, but we’ve doubled their ranks: we now have over 50 leading computer scientists — one in three of which is a unicorn founder, one in five a decacorn founder. This includes heroes of mine like Turing Award winner Dave Patterson and Jeff Dean, two of my co-authors on a recent <a href="http://shapingai.com">AI vision paper</a>. In addition to doubling our research LP base, Andrew and I have joined forces with Pete Sonsini, an 18-year veteran that built his career investing in researcher founders including me and my Databricks and Perplexity co-founders.</p>

<p>I’m obsessed with researchers that have changed the world through open source and through startups, and I’m lucky to be one of them. I can’t wait to inspire, support, and fund ten more companies that move us all forward, just like Databricks and Perplexity have.</p>]]></content><author><name>Andy Konwinski</name></author><summary type="html"><![CDATA[A few years ago, I co-founded a small fund called CSGV (CS Grad Ventures) that took money from — and invested it in — Computer Science PhDs and Professors. Our vision was to democratize Venture Capital back to the researchers themselves. To succeed we would need to find the next Databricks.]]></summary></entry><entry><title type="html">Introducing Headless Terminal</title><link href="/2024/05/31/introducing-headless-terminal.html" rel="alternate" type="text/html" title="Introducing Headless Terminal" /><published>2024-05-31T00:00:00+00:00</published><updated>2024-05-31T00:00:00+00:00</updated><id>/2024/05/31/introducing-headless-terminal</id><content type="html" xml:base="/2024/05/31/introducing-headless-terminal.html"><![CDATA[<p>Introducing <em><a href="https://github.com/andyk/ht">Headless Terminal (<code class="language-plaintext highlighter-rouge">ht</code>)</a></em> - making terminals easy for LLMs to use.</p>

<p>I’ve been using LLM agents for coding, and needed something like a headless browser but for terminals. So I teamed up with <a href="https://x.com/sickill">@sickill</a> (creator of <a href="https://asciinema.org">asciinema</a>) to build <a href="https://github.com/andyk/ht"><code class="language-plaintext highlighter-rouge">ht</code></a>.</p>

<p>Headless Terminal is an open source executable that wraps and provides text screenshots of a terminal. Terminals are one of the oldest and most prolific UI frameworks in all of computing. And they are stateful so, for example, when you use an editor in your terminal, the terminal has to manage state about the cursor location. Without <code class="language-plaintext highlighter-rouge">ht</code>, an agent struggles to manage this state directly. With <code class="language-plaintext highlighter-rouge">ht</code>, an agent can just observe the terminal like a human does.</p>

<p>How <code class="language-plaintext highlighter-rouge">ht</code> works:</p>

<ul>
  <li>The <code class="language-plaintext highlighter-rouge">ht</code> binary wraps an arbitrary other binary (e.g. bash, vim, etc.) with a VT100 style terminal interface.</li>
  <li>By default, <code class="language-plaintext highlighter-rouge">ht</code> starts bash, but you can override it to wrap any binary (e.g., <code class="language-plaintext highlighter-rouge">ht nano</code>)</li>
  <li>Communication with <code class="language-plaintext highlighter-rouge">ht</code> is performed via stdin, stdout and stderr.</li>
  <li><code class="language-plaintext highlighter-rouge">ht</code> uses simple JSON-based protocol for sending input to its stdin and fetching screenshots of the terminal.
    <ul>
      <li>Input is sent as a JSON object in the form { “type”: “input”, “payload”: “ls\r” }</li>
      <li>The agent can “look at” the terminal, i.e., grab a screenshot, by sending the following to JSON object: { “type”: “getView” }</li>
    </ul>
  </li>
  <li>Diagnostic messages (notices, errors) are printed to stderr.</li>
</ul>

<p><code class="language-plaintext highlighter-rouge">ht</code> is Apache licensed, built in rust, and works on MacOS and Linux.</p>

<center><img src="/assets/img/headless-terminal.png" style="width: 85%" /></center>]]></content><author><name>Andy Konwinski</name></author><summary type="html"><![CDATA[Introducing Headless Terminal (ht) - making terminals easy for LLMs to use.]]></summary></entry><entry><title type="html">LLMs are like toddlers</title><link href="/2024/05/14/llms-are-like-toddlers.html" rel="alternate" type="text/html" title="LLMs are like toddlers" /><published>2024-05-14T00:00:00+00:00</published><updated>2024-05-14T00:00:00+00:00</updated><id>/2024/05/14/llms-are-like-toddlers</id><content type="html" xml:base="/2024/05/14/llms-are-like-toddlers.html"><![CDATA[<center><a href="https://twitter.com/andykonwinski/status/1790485847013523758"><img alt="Tweet: Can't stop seeing similarities between my 3yr old and Llama 3. Few shot learning, wild hallucinations, insane cost to create. Can we make training models as inexplicably satisfying as explaining the plot of Frozen to a toddler… for the 100th time?" src="/assets/img/llm-toddler-tweet-05-14-2024.png" style="width:500px; padding-bottom:30px" /></a></center>

<p>I think my 3yr old and Llama 3 were separated at birth. The way training my daughter involves (verbally) collecting a dataset of thousands of examples. The way she sometimes needs only one example, sometimes dozens. The way she hallucinates (e.g., while learning to count, she skipped “13” for months).</p>

<p>However, thanks to my biological imperitive as a dad, I find it deeply satisfying to spend hundreds of hours generating a bespoke training dataset for her in the form of me continuously explaining how the world works (e.g., today we covered the difference between smoke and steam). That dataset will never be written down and can never be used to train another neural network besides my 1 year old overhearing it all. Whereas for LLMs, the training set is larger, noisier, mostly static, and I don’t feel the same emotional drive to improve it.</p>

<p>Maybe we should hack our biology by making LLMs look or behave more like small children. Or maybe I should sprinkle Raspberry Pi’s all over my house that transcribe all of our conversations into a dataset for pre-training?</p>]]></content><author><name>Andy Konwinski</name></author><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">AI benchmarks should be like unit tests</title><link href="/2024/05/08/ai-benchmarks-should-be-like-unit-tests.html" rel="alternate" type="text/html" title="AI benchmarks should be like unit tests" /><published>2024-05-08T00:00:00+00:00</published><updated>2024-05-08T00:00:00+00:00</updated><id>/2024/05/08/ai-benchmarks-should-be-like-unit-tests</id><content type="html" xml:base="/2024/05/08/ai-benchmarks-should-be-like-unit-tests.html"><![CDATA[<p>Today’s go-to LLM benchmarks incentivize model makers to improve marginal
performance on tasks that AIs (i.e., a model or model+memory+tools+etc.) have
already mastered at superhuman-levels like general Q&amp;A, fact recall,
summarization, simple few-step reasoning, etc. In contrast, my favorite
benchmarks (like <a href="https://swebench.com">SWE-bench</a>) are more like unit tests in
that they test a skill which AIs still haven’t achieved at human-level, e.g.,
complex logical inference, tool usage, long term memory, knowing what they
don’t know, integrating actions + observations + reasoning, learning how to
learn.</p>

<p>By design, today’s best AIs will perform miserably on such a benchmark
(analogous to a unit test failing), which will focus research and engineering
efforts on high-leverage areas of improvement.</p>

<p><a href="https://swebench.com">SWE-bench</a> is a good example. #1 on their leaderboard is
currently at 12.5%, and #2 is at than 3.8%. Can we build benchmarks like this for
use-cases beyond software engineering? Maybe healthcare, finance, data
science/eng., personal assistant, social influencer…?</p>]]></content><author><name>Andy Konwinski</name></author><summary type="html"><![CDATA[Today’s go-to LLM benchmarks incentivize model makers to improve marginal performance on tasks that AIs (i.e., a model or model+memory+tools+etc.) have already mastered at superhuman-levels like general Q&amp;A, fact recall, summarization, simple few-step reasoning, etc. In contrast, my favorite benchmarks (like SWE-bench) are more like unit tests in that they test a skill which AIs still haven’t achieved at human-level, e.g., complex logical inference, tool usage, long term memory, knowing what they don’t know, integrating actions + observations + reasoning, learning how to learn.]]></summary></entry><entry><title type="html">Focus on the failures</title><link href="/2023/09/18/focus-on-the-failures.html" rel="alternate" type="text/html" title="Focus on the failures" /><published>2023-09-18T00:00:00+00:00</published><updated>2023-09-18T00:00:00+00:00</updated><id>/2023/09/18/focus-on-the-failures</id><content type="html" xml:base="/2023/09/18/focus-on-the-failures.html"><![CDATA[<p>When playing LLMs, it is tempting to only share and discuss the prompts that cause the model to return something interesting that demonstrates the models reasoning capabilities or the knowledge it has memorized.</p>

<p>For me, sometimes those prompts come after a lot of iteration. Less often, they are the first thing I write after the general idea pops into my head.</p>

<p>Most of the lessons in prompt design, and more generally systems design, come from the iterative process. The failures are just as important as the successes. They are the things that didn’t work, and the things that didn’t work are the things that we can learn from.</p>]]></content><author><name>Andy Konwinski</name></author><summary type="html"><![CDATA[When playing LLMs, it is tempting to only share and discuss the prompts that cause the model to return something interesting that demonstrates the models reasoning capabilities or the knowledge it has memorized.]]></summary></entry><entry><title type="html">List of Autonomous LLM Agent Projects</title><link href="/2023/03/30/list-of-autonomous-agent-work.html" rel="alternate" type="text/html" title="List of Autonomous LLM Agent Projects" /><published>2023-03-30T00:00:00+00:00</published><updated>2023-03-30T00:00:00+00:00</updated><id>/2023/03/30/list-of-autonomous-agent-work</id><content type="html" xml:base="/2023/03/30/list-of-autonomous-agent-work.html"><![CDATA[<p><em>Last Updated: April 19, 2023</em></p>

<p>This is my scratchpad of interesting projects and products focused on building agents on top of LLMs that (1) can do things via actions &amp; tools, (2) autonomously decide their own goals. There is also a short summary of each project.</p>

<ul>
  <li><a href="https://github.com/Torantulino/Auto-GPT">Auto-GPT</a>, Toran Bruce Richards</li>
  <li><a href="https://arxiv.org/abs/2302.04761">Toolformer</a>, Timo Schick et al., Meta AI Research, Universitat Pompeu Fabra</li>
  <li><a href="https://ai.googleblog.com/2022/11/react-synergizing-reasoning-and-acting.html">ReAct - Reasoning + Actions</a>, Shunyu Yao et al., Google Brain</li>
  <li><a href="https://github.com/yoheinakajima/babyagi">BabyAGI</a>, Yohei Nakajima</li>
  <li><a href="https://www.ai21.com/blog/jurassic-x-crossing-the-neuro-symbolic-chasm-with-the-mrkl-system">MERKL - Modular Reasoning, Knowledge and Language</a>, AI21</li>
  <li><a href="https://arxiv.org/abs/2211.10435">PAL - Program Aided Language Models</a>, Luyu Gao et al., CMU</li>
  <li><a href="https://python.langchain.com/en/latest/modules/agents.html">LangChain Agents</a>, Harrison Chase</li>
  <li><a href="https://openai.com/blog/chatgpt-plugins">GPT Plugins</a>, OpenAI
    <ul>
      <li><a href="https://www.getit.ai/gpt-plugins">getit.ai plugins &amp; agent registry</a></li>
    </ul>
  </li>
  <li><a href="https://selfrefine.info/">Self-Refine: Iterative Refinement with Self-Feedback</a> (<a href="https://arxiv.org/abs/2303.17651">paper</a>), Aman Madaan et al., CMU, Allen Institute, U of Washington, etc.</li>
  <li><a href="https://dust.tt/">Dust</a>, Stanislas Polu</li>
  <li><a href="https://github.com/stanfordnlp/dsp/">DSP - Demonstrate Search Predict</a>, Omar Khattab et al., Stanford</li>
  <li><a href="https://arxiv.org/abs/2303.11366">Reflexion</a>, Northeastern, MIT</li>
  <li><a href="https://www.toolkit.club/">toolkit.club</a> - LLM plugin registry</li>
</ul>

<p><strong>Related reads (no code or project per se):</strong></p>
<ul>
  <li><a href="https://arxiv.org/abs/2205.11916">Large Language Models are Zero-Shot Reasoners</a> - the “let’s think step by step” paper, Kojima et al., The University of Tokyo, Google</li>
  <li><a href="https://jmcdonnell.substack.com/p/the-near-future-of-ai-is-action-driven?sd=pf">The Near Future of AI is Action-Driven</a>, James McDonnell</li>
  <li><a href="https://arxiv.org/pdf/2303.12712.pdf">Sparks of Artificial General Intelligence Early experiments with GPT-4</a> Section 5: “Interaction with the world” (page 43), Seb́astien Bubeck et al., Microsoft</li>
  <li><a href="https://arxiv.org/abs/2304.03442">Generative Agents: Interactive Simulacra of Human Behavior</a>, Park et al., Stanford, Google</li>
  <li><a href="https://twitter.com/karpathy/status/1642598890573819905">Tweet thread by Andrej Karpathy re “AutoGPTs”</a></li>
  <li><a href="https://twitter.com/mathemagic1an/status/1645096275392745477">Tweet thread by Jay Hack</a> on auto agents.</li>
</ul>

<p>Most of the summaries below were generated by the <a href="https://perplexity.ai">perplexity.ai</a> <a href="https://chrome.google.com/webstore/detail/perplexity-ask-ai/hlgbcneanomplepojfcnclggenpcoldo">Chrome extension</a>.</p>

<h3 id="toolformer">Toolformer</h3>

<p><strong><a href="https://arxiv.org/pdf/2302.04761.pdf">Toolformer: Language Models Can Teach Themselves to Use Tools</a></strong>, Timo Schick et al., Meta AI Research, Universitat Pompeu Fabra</p>

<p>The paper introduces a new fine-tuned model called Toolformer, which is trained to use external tools via simple APIs in a self-supervised way. The model decides which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction. The model incorporates a range of tools, including a calculator, a Q&amp;A system, a search engine, a translation system, and a calendar. Toolformer achieves substantially improved zero-shot performance across a variety of downstream tasks, often competitive with much larger models, without sacrificing its core language modeling abilities.</p>

<h3 id="react">ReAct</h3>

<p><strong><a href="https://arxiv.org/pdf/2210.03629.pdf">ReAct: Synergizing Reasoning and Acting in Language Models</a></strong>, Shunyu Yao et al., Google Brain</p>

<p>The paper discusses the use of large language models (LLMs) to generate both reasoning traces and task-specific actions in an interleaved manner, allowing for greater synergy between the two. The authors apply their approach, named ReAct (for Reasoning + Action), to a diverse set of language and decision-making tasks and demonstrate its effectiveness over state-of-the-art baselines in addition to improved human interpretability and trustworthiness. Specifically, on question answering and fact verification tasks, ReAct overcomes prevalent issues of hallucination and error propagation in chain-of-thought reasoning by interacting with a simple Wikipedia API, and generating human-like task-solving trajectories that are more interpretable than baselines without reasoning traces. Furthermore, on two interactive decision-making benchmarks, ReAct outperforms imitation and reinforcement learning methods by a significant margin.</p>

<h3 id="babyagi">BabyAGI</h3>
<p><strong><a href="https://github.com/yoheinakajima/babyagi">BabyAGI</a></strong>, Yohei Nakajima
Twitter threads with the <a href="https://twitter.com/yoheinakajima/status/1642881722495954945">original announcement</a> and the <a href="https://twitter.com/yoheinakajima/status/1640934493489070080">open source release</a>.</p>

<p>Copied from the github README:</p>

<blockquote>
  <p>This Python script is an example of an AI-powered task management system. The system uses OpenAI and Pinecone APIs to create, prioritize, and execute tasks. The main idea behind this system is that it creates tasks based on the result of previous tasks and a predefined objective. The script then uses OpenAI’s natural language processing (NLP) capabilities to create new tasks based on the objective, and Pinecone to store and retrieve task results for context. …</p>

  <p><strong>How It Works</strong></p>

  <p>The script works by running an infinite loop that does the following steps:</p>

  <ol>
    <li>Pulls the first task from the task list.</li>
    <li>Sends the task to the execution agent, which uses OpenAI’s API to complete the task based on the context.</li>
    <li>Enriches the result and stores it in Pinecone.</li>
    <li>Creates new tasks and reprioritizes the task list based on the objective and the result of the previous task. The execution_agent() function is where the OpenAI API is used. It takes two parameters: the objective and the task. It then sends a prompt to OpenAI’s API, which returns the result of the task. The prompt consists of a description of the AI system’s task, the objective, and the task itself. The result is then returned as a string.</li>
  </ol>
</blockquote>

<h3 id="merkl">MERKL</h3>
<p><strong><a href="https://www.ai21.com/blog/jurassic-x-crossing-the-neuro-symbolic-chasm-with-the-mrkl-system">Jurassic-X: Crossing the neuro-symbolic chasm with the MRKL system</a></strong>, AI21</p>

<p>MRKL (Modular Reasoning, Knowledge and Language) is a system designed to bridge the gap between symbolic reasoning and neural networks. The system uses a DNN to classify incoming messages and creating a plan for a series of calls to expert modules, for examples extracting information from multiple sources and summarizing them. They have built modules for using a database, a calculator, etc. AI21 has a proprietry implementation of the MRKL system, called Jurassic-X, which is being piloted by a few partners.</p>

<h3 id="pal-program-aided-language-models">PAL: Program-aided Language Models</h3>

<p><strong><a href="https://arxiv.org/pdf/2211.10435.pdf">PAL: Program-aided Language Models</a></strong>, Luyu Gao et al., CMU</p>

<p>The paper presents a new approach called Program-Aided Language models (PAL) that uses large language models (LLMs) to read natural language problems and generate programs as intermediate reasoning steps, but delegates the solution step to a runtime such as a Python interpreter. The authors demonstrate that this approach leads to more accurate results than other methods, including chain-of-thought, on 13 mathematical, symbolic, and algorithmic reasoning tasks. The paper argues that while LLMs are adept at step-by-step decomposition, they often make logical and arithmetic mistakes in the solution part. PAL addresses this issue by using the LLM to generate programs and offloading the solution step to an interpreter. The authors provide code and data publicly.</p>

<h3 id="langchain-agents">LangChain Agents</h3>

<p><strong><a href="https://python.langchain.com/en/latest/modules/agents/getting_started.html">LangChain Agents docs</a></strong>, Harrison Chase &amp; Open Source contributors. The summary below was copied and tweaked from one of <a href="https://twitter.com/hwchase17/status/1595456660507459585">Harrison’s tweet threads</a>.</p>

<p>In Langchain, Agents use an LLM to determine which which actions to take and in what order. An action can either be using a tool and observing its output, or returning to the user. “Tools” in this context can be anything that takes a string as input and outputs a string. This can be a search engine, a database, another LLM, a chain, or event another agent.</p>

<p>The abstractions in Langchain were inspired by much of the above work. In the tweet thread linked above, Harrison Chase said “This is NOT a new idea, just a reframing”:</p>

<blockquote>
  <p>I would argue that @ShunyuYao12’s ReAct paper (<a href="https://arxiv.org/pdf/2210.03629.pdf">https://arxiv.org/pdf/2210.03629.pdf</a>) uses an agent, which has a “Search” tool and a “Lookup” tool available</p>
</blockquote>

<blockquote>
  <p>Likewise, I would argue that @OfirPress’s self-ask paper (<a href="https://arxiv.org/abs/2210.03350">https://arxiv.org/abs/2210.03350</a>) uses an agent, which has a “Search” tool available</p>
</blockquote>

<blockquote>
  <p>The MRKL implementation (inspired by @AI21Labs <a href="https://www.ai21.com/blog/jurassic-x-crossing-the-neuro-symbolic-chasm-with-the-mrkl-system">ai21.com/blog/jurassic-x…</a>) is extremely close to this current agent abstraction</p>
</blockquote>

<h3 id="gpt-plugins">GPT Plugins</h3>

<p><strong><a href="https://openai.com/blog/chatgpt-plugins">GPT Plugins</a></strong>, OpenAI</p>

<p>Plugins aim to be the “eyes and ears” for language models, giving them access to information that is too recent, too personal, or too specific to be included in the LLM’s training data. They can also enable language models to perform actions on behalf of users.</p>

<p>Plugin developers expose one or more API endpoints, accompanied by a standardized manifest file (hosted at <code class="language-plaintext highlighter-rouge">yourdomain.com/.well-known/ai-plugin.json</code>) and an <a href="https://swagger.io/specification/">OpenAPI specification</a>. Then ChatGPT (or some other AI model) acts as an intelligent API caller, given an API spec and a natural-language description of when to use the API. To build a plugin, developers need to create a manifest file, register the plugin in the ChatGPT UI, and have users activate the plugin. When a user asks a relevant question, the model may choose to invoke an API call from the plugin if it seems relevant, and incorporate the API results into its response to the user.</p>

<p>The features were initially relased as a private beta that required users to joint a waitlist. The first plugins were created by a handful of popular web companies including as Slack, Wolfram, and Zapier. OpenAI is also hosting two plugins themselves: a web browser and code interpreter. Besides the initial plugins, there is also a lugin API that allows anyone to create their own plugins.</p>

<h2 id="dust">Dust</h2>

<p><strong><a href="https://github.com/dust-tt/dust">Dust</a></strong> (<a href="https://dust.tt/">hosted version</a>) (<a href="https://docs.dust.tt/overview">docs</a>), Stanislas Polu &amp; open source contributors</p>

<p>The Dust Platform is a framework (written in Rust) designed to define and deploy large language model apps without writing any execution code. It aims to ease the process of working on multiple examples at the same time, introspecting model outputs produced by intermediary steps of large language model apps, and iterating on the design of large language model apps by providing a granular and automated versioning system. Dust apps are composed of blocks executed sequentially, and each block has a set of specification arguments and configuration arguments. The input block is the block that receives the arguments required to run a Dust app, and Dust apps can also define datasets which are arrays of JSON objects. The Dust execution engine will run an app in parallel on each element attached to the input block.</p>

<h3 id="dsp">DSP</h3>

<p><strong><a href="https://github.com/stanfordnlp/dsp/">Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP</a></strong> (<a href="https://arxiv.org/abs/2212.14024">paper</a>)</p>

<p>The Demonstrate-Search-Predict (DSP) framework is a programming abstraction for building AI systems for knowledge-intensive tasks such as answering user questions or researching complex topics. DSP programs are written in a few lines of code, describing how the problem should be decomposed into smaller transformations. The framework provides primitives for composing transformations and and mapping these transformations to effective language model (LM, GPT-3.5 in this case) and retrieval model (RM - ColBERTv2) calls. DSP discourages “prompt engineering” and offers a number of automatic tuning features. The evaluation in the paper shows DSP beat vanilla GPT-3.5 on Open-SQuAD by 120% (i.e., DSP got a score of 36.6 vs. vanilla GPT-3.5 with a few-shot prompt’s score of 16.2). They beat the simpler Retreive-then-read paradigm by a more modest 8%. They also evaluate performance on HotPotQA, and QReCC. The framework is open source and available for installation and comes with an intro notebook and compiler notebook.</p>

<h3 id="reflexion">Reflexion</h3>

<p><strong><a href="https://arxiv.org/abs/2303.11366">Reflexion: an autonomous agent with dynamic memory and self-reflection</a></strong></p>

<p>Snippet from the abstract: “Building on recent research, we propose Reflexion, an approach that endows an agent with dynamic memory and self-reflection capabilities to enhance its existing reasoning trace and task-specific action choice abilities. To achieve full automation, we introduce a straightforward yet effective heuristic that enables the agent to pinpoint hallucination instances, avoid repetition in action sequences, and, in some environments, construct an internal memory map of the given environment. To assess our approach, we evaluate the agent’s ability to complete decision-making tasks in AlfWorld environments and knowledge-intensive, search-based question-and-answer tasks in HotPotQA environments. We observe success rates of 97% and 51%, respectively.”</p>

<h3 id="end-note">End note</h3>

<p>Email me (or create a <a href="https://github.com/andyk/andyk.github.io/blob/master/_posts/2023-03-30-list-of-autonomous-agent-work.md">PR</a>) with anything that I’m missing and I’ll update this page.</p>]]></content><author><name>Andy Konwinski</name></author><summary type="html"><![CDATA[Last Updated: April 19, 2023]]></summary></entry></feed>