Brian Bell · Oct 8, 2026 · 9 min read

Ignite AI: The AI Gateway for Models, Agents and Multi-Device Inference with Jonathan Archer | Ep301

Jonathan Archer of Open LLM explains how to get more from your AI stack by combining multiple models, subscriptions, agents, and devices without constantly switching tools.

Ignite AI: The AI Gateway for Models, Agents and Multi-Device Inference with Jonathan Archer | Ep301

AI users are increasingly paying for several models at once, but the experience is still fragmented.

One model may be better at writing. Another may be better at image generation. Another may have stronger access to social data. Developers may also have API keys, local models, cloud instances, and agents running across multiple machines.

Jonathan Archer, co-founder and CEO of Open LLM, believes the answer is not choosing one winner. It is building infrastructure that lets all of those models work together.

In his conversation with Brian Bell on the Ignite Podcast, Jonathan explains why he expects AI to remain a multi-model ecosystem, why users should retain control over their subscriptions and tokens, and why the gateway layer may evolve into something much larger than simple model routing.

Open LLM Started With Rate Limits

The original problem was straightforward.

Jonathan’s co-founder was repeatedly hitting rate limits while coding. When that happened, he began experimenting with open-source alternatives that could combine access across different providers.

The problem was that maintaining those systems became work in itself.

Jonathan experienced the same friction. The alternatives were possible to configure, but they were complicated enough that the setup and maintenance often outweighed their value.

That led to the idea behind Open LLM: create a gateway specifically designed to bring AI subscriptions and API access together.

Jonathan argues that subscriptions can currently offer unusually strong economics. In the episode, he says a $200 Claude subscription can provide almost $10,000 worth of usage, although he expects the amount of usage included in those plans to decline over time.

If that happens, bundling multiple subscriptions becomes increasingly useful.

The Future Is Multi-Model

One of Jonathan’s strongest convictions is that there will not be one AI model that handles every task best.

He points to several practical differences already emerging between providers.

Grok can be useful for searching Twitter. ChatGPT can be used for image generation. Claude, in Jonathan and Brian’s view, performs particularly well for writing.

That creates a simple problem for users.

If the best workflow requires three or four different models, constantly jumping between interfaces becomes inefficient.

Open LLM is designed around the opposite approach. Users bring their own subscriptions and API keys, then use those resources from the environment where they already prefer to work.

Jonathan summarizes the thesis simply:

“We’re gonna live in a multi-model future.”

The goal is not to force users into another closed interface. It is to let them use different models without constantly changing where they work.

Your Tokens Should Not Be Locked Into One Harness

Jonathan is particularly critical of platforms that try to lock users into their own AI environment.

His argument is that customers are already paying for access to the models. They should be able to use that access wherever they want.

For Open LLM, that means supporting subscriptions and API keys while allowing users to continue working from tools such as Claude Code, Codex, or other environments.

Jonathan’s view is:

“It’s their tokens, they should be able to go.”

That philosophy shapes the product.

Rather than trying to become the only AI interface a customer uses, Open LLM aims to become infrastructure underneath the tools they already use.

Multi-Device Inference Could Become Much More Important

One of the more interesting parts of the conversation is Open LLM’s approach to devices.

Jonathan describes running multiple machines, including virtual private servers, while only some of those machines are directly authenticated into his AI subscriptions.

Instead of forcing every new machine to authenticate again, another device can pull inference from a machine that already has access.

That becomes more significant when thinking about autonomous agents.

Jonathan expects agents to increasingly spin up machines and create subagents to complete individual tasks.

Those agents will still need inference.

They could rely on API keys, or they could potentially access inference through another machine already connected to the user’s subscriptions.

The same architecture could also support local models running on stronger hardware elsewhere.

For Jonathan, this is one reason Open LLM is moving beyond the idea of being merely a gateway.

The Right Model for the Right Job

Open LLM is not currently positioned as a fully autonomous smart router that perfectly determines which model should handle every request.

Jonathan is careful about that distinction.

Instead, the platform currently supports configurable fallback chains.

A user can organize models into different intelligence tiers. If one model hits its usage limit or its provider experiences downtime, Open LLM can move to another model in the same tier.

The system can also differentiate between larger primary tasks and smaller subagent tasks.

Jonathan argues that using the most capable model for every small action is wasteful.

His analogy is simple: you would not send Steve Jobs to retrieve a file.

The same principle applies to AI infrastructure.

High-intelligence models should handle tasks that justify their cost and capability, while smaller models can handle routine subagent work.

Open LLM Is Only Partly Open Source

Open LLM also takes a hybrid approach to open source.

The background process running on the user’s computer is open so that users can inspect it and verify what it is doing.

The frontend and dashboard remain closed.

Jonathan says the dashboard simplifies setup and gives users visibility into their system, including chat and other management features.

The company’s differentiation, however, is not simply open versus closed source.

Jonathan says the team wants to stay extremely responsive to the community and increasingly operate in an agent-first way.

He even discusses the possibility of an AI interacting directly with users, collecting product requests, starting work on the requested code, and then submitting it for human review.

From Gateway to Complete AI Infrastructure

The long-term ambition is broader than routing requests between models.

Jonathan says Open LLM wants to become “all your AI infra in one place.”

That roadmap includes multi-device sessions, shared inference, unified memory, tunneling between machines, and potentially one-click deployment of virtual private servers.

One planned use case illustrates the direction.

A user could be on a train with only a phone, connect through the Open LLM interface to another computer, access a coding agent running there, and begin shipping code without opening a laptop.

For enterprise customers, the architecture changes because subscriptions are not appropriate at that scale. Jonathan says those organizations would use API keys instead.

The company is also thinking about shared memory and indexed codebases across teams.

The broader thesis is that AI infrastructure will increasingly involve models, agents, devices, memory, and compute working together rather than sitting inside one chatbot window.

Dynamic Discovery and a Faster AI Stack

Another feature Jonathan describes is dynamic discovery.

Open LLM continuously queries model providers to see what models are available.

He says that can occasionally result in a new model appearing inside Open LLM before the provider publicly announces it.

The system is also being designed so that users can modify parts of their AI infrastructure directly through a command-line interface.

Instead of opening a dashboard and manually changing a fallback chain, Jonathan wants users to be able to simply ask an AI tool to make the change.

That reflects a broader belief about software.

Traditional cloud infrastructure can be powerful but overwhelming. Jonathan contrasts that with newer tools that allow developers to manage infrastructure directly from their working environment.

The interface itself increasingly becomes conversational.

Crypto May Have Found a Native AI Use Case

The episode eventually circles back to Jonathan’s background in crypto.

He says he once expected technologies such as NFTs to become widely used for things like passports, IDs, and deeds.

The technology may support those applications, but adoption has proven much harder.

Where Jonathan still sees obvious practical value is in moving money.

He points to stablecoins and crypto adoption in places such as Argentina, where inflation creates stronger demand for alternatives to holding local currency.

But his more unusual claim is about AI agents.

Jonathan suggests that crypto may ultimately prove especially useful for agents rather than humans.

An agent can create a wallet and transact programmatically without needing the same banking interfaces that humans use.

As autonomous agents become capable of acquiring compute, services, data, and other resources, native digital payment rails may become increasingly valuable.

The Gateway Layer May Survive

Brian asks whether AI gateways will remain standalone companies five years from now or eventually disappear into cloud providers, model companies, and agent frameworks.

Jonathan believes the category can remain independent, but only if it expands.

Routing alone is not enough.

That is why Open LLM is pushing toward multi-device management, shared inference, unified memory, and broader AI infrastructure.

Jonathan also says there is a strong possibility that Open LLM could eventually be acquired by a larger infrastructure company.

But his immediate focus is straightforward.

If everything disappeared tomorrow, he says he would rebuild Open LLM.

He believes there is still enough time to carve out an important position in the AI infrastructure stack.

The Bigger Shift

The most important idea from the conversation is not that one model is better than another.

It is that model choice itself may become less important to the end user.

If AI infrastructure can reliably route work across different models, subscriptions, devices, and agents, the user may eventually care less about which provider completed the task.

The winning layer could be the one that makes all of those resources behave like a single system.

That is the future Jonathan Archer is betting Open LLM can help build.

Chapters:
00:01 Jonathan Archer and Open LLM

  • 03:34The Origin of Open LLM
  • 05:30AI Subscription Economics
  • 07:28Subscription Routing and Compliance
  • 09:01Unified AI Memory
  • 11:04Multi-Device AI Infrastructure
  • 13:19Self-Hosted Models and VPS Deployment
  • 14:43Open LLM Pricing and Free Tier
  • 16:19Cross-Model Workflows
  • 17:03The Technical Challenge of AI Routing
  • 19:10The Multi-Model AI Future
  • 20:24Free vs Paid Plans
  • 21:37Fallback Chains and Model Selection
  • 23:38AI Subagents and Token Efficiency
  • 24:35Open Source vs Closed Infrastructure
  • 26:12Multi-Device Sessions and Enterprise AI
  • 28:15Breaking AI Platform Lock-In
  • 29:53OpenRouter and the Gateway Model
  • 32:06AI Gateway Acquisition Value
  • 34:09Open LLM and Vercel
  • 36:18Dynamic Model Discovery
  • 37:38Beyond the AI Gateway
  • 38:29Rapid-Fire Questions
  • 39:00Crypto Adoption and Real-World Utility
  • 40:56Crypto, Payments, and Stablecoins

Listen to this episode

0:00 / 0:00
Open the full episode page ↗
Read the full transcript

Brian Bell (00:00.817) Hey everyone, welcome back to the Ignite Podcast. Today we are delighted to have Jonathan Archer on the mic. He's the co founder and CEO of Open LLM, a gateway that puts every AI model you pay for, Cloud, GPT, Gemini, et cetera, behind one key. Before building infrastructure for the AI stack, he spent years in web three running community growth and tokenics for NFT projects. He's here to talk about what that takes to build an AI right now. Thanks for coming on, Jonathan.

Jonathan Archer (00:25.974) Yeah, thanks so much for having me, Brian. I'm stoked to be here.

Brian Bell (00:28.273) Well I'd love to s yeah. I'd love to start with your origin story. What's your background? Besides what I said.

Jonathan Archer (00:33.378) Ye yeah, so as you kinda mentioned, you know, I've been advising and helping out with a bunch of different crypto projects for the past six years or so. And that's actually how I met my co-founder, Mamet. we met at a crypto conference in Buenos Aires, the Avalanche Summit, about two years ago. And it was one of those funny connections where you meet someone and just immediately hit it off. So him and I started kinda trying to cook up some fun side projects online until eventually we started working together in person. but yeah, really when it comes to B D growth, community management, everything in the non-technical side of project development.

Brian Bell (01:15.557) Yeah, and we were we were joking in in front of the podcast that you did a a teacher training program for yoga. What what was that that like and how does that inform your kind of life as a founder?

Jonathan Archer (01:26.518) Yeah, so the teacher training crogram was amazing, but what really taught me a lot was actually my time meditating. So I lived in India for about a year and a half, just meditating almost every day, spending a lot of time in silent retreats, and I think

Brian Bell (01:42.779) Was this in the Vedic tradition, so you like mantras and stuff like that?

Jonathan Archer (01:46.423) I did a bit of that, but I'm actually more in line with the Buddhist traditions. And the last school I practiced was the was really focused on the Theravadas. So like the the the closest thing to the sutras or what the Buddha directly taught. more let's say yeah, more traditional. And I really enjoyed that because

Brian Bell (01:50.256) Okay, gotcha.

Brian Bell (02:08.933) Yeah. More like in insights of things in your mind, breath, body sensations, stuff like that.

Jonathan Archer (02:14.814) Exactly. And actually just open awareness, you know, and bringing that calm mind to whatever it is where I've been in situations with my co-founder that can be very, very stressful. And I've noticed that this meditation practice has actually helped me maintain that calm, you know, and I I think that's really important and and one of the best skills that you can nurture as a founder is like, okay, you know, you have to be able to push yourself and like

Brian Bell (02:17.659) Yeah. Open awareness, yeah.

Jonathan Archer (02:43.16) Take that stress and use it in a positive way, but you can't let it overcome you and just take advantage of the situation.

Brian Bell (02:51.697) Yeah. Yeah, I agree. I'm a naturally kind of a more hot headed person, believe it or not. And I think the meditation kind of smooths out the rough edges a little bit, right?

Jonathan Archer (03:01.504) Yeah, and it's so easy to be hard on oneself, right? And I think that's something that my meditation teacher was always always telling me, he's like, Jonathan, you being hard on yourself again and you know, you just gotta loosen up and have fun and when you're doing that, everyone has more fun around you, there's better deals, you know, things just flow naturally.

Brian Bell (03:20.539) Yeah. Yeah. So you guys meet at a crypto conference of sorts and at some point you guys make this leap into AI. What was that moment behind you know, every model behind one key for you?

Jonathan Archer (03:34.403) So, you know, my co-founder will tell you that I'm the guy who's got way too many ideas. but this is the one idea that I can take credit for. It was actually from his experience of getting rate limited all the time. And I would also get rate limited when using these subscriptions, but he was getting rate limited like every day, you know, and that's just from coding. You know, coders naturally are gonna use a lot more tokens than the marketing, you know, BD, whatever it is, kind of guys. and so when he did hit those limits, he started looking for open source alternatives, you know, okay, what's out there? What can he start doing to actually start bundling subscriptions together? And he found a few, but unfortunately he found that he was spending more time actually maintaining them than getting value from them. And to set it up, it's a big drag. For a non-technical user, it's quite complicated. and I even set them up myself and I was able to, but I felt similar pains. And so once we realized that, we said, okay, this is a real problem. If he's facing it, I'm facing it not quite as much, but still enough to be interested in a product like this, then we should actually develop it. And so that's when we decided, all right, you know, a gateway for subscription specifically and APIs makes a lot of sense.

Brian Bell (04:57.169) Yeah, I I face this problem myself, you know, because Team Ignite runs on AI. And you know, I think I'm on like the super max cloud, whatever the three hundred dollar a month subscription is, to get a lot of usage because I have constant agents running all the time. And every once in a while I hit the w or the the rate limit. Not as much anymore on the three hundred dollar a month plan, but you know, and I'll I'll use Chat GPT as a backup. So I'm kind of manually doing that. But I can totally empathize with with coders who who need They need the code, right? They need the code right now and and they're using a lot of tokens to get it.

Jonathan Archer (05:30.828) Yeah, and so you know, obviously using the subscription saves you a lot of money. So these subscriptions are giving you twenty, thirty, forty, fifty X. Like with a two hundred dollar plan from Claude, you're getting almost ten thousand dollars worth of usage, believe it or not. our thesis is a few different things. One, we're gonna see that number start to drop. So it's gonna drop from ten K to eight to six, you know, to five. And when it does

Brian Bell (05:57.554) They they know what they're doing. They're trying to they're kind of doing the price elasticity calculations. They're not stupid. They probably they have super intelligent AI internally and they're like, okay, like, you know, please maximize my revenue, you know, per hour of usage or whatever. And they're like, if we just kind of like tier it this way and price discriminate that way and we can, you know, get the maximum number of tokens. I think about this a lot when I ask the AI to query something. I'm like, you used a lot of tokens for something pretty simple, you know.

Jonathan Archer (06:26.707) Right. Right. And you're like, what's going on?

Brian Bell (06:27.813) You know, and I'm like, that makes sense. They're trying to maximize usage so I can pay for the max plan or whatever.

Jonathan Archer (06:33.976) Well also it looks amazingly bullish to investors, you know. It's like look how many tokens have gone through. And I I saw a funny meme meme on LinkedIn the other day, which I I never thought those words would come out of my mouth, but they they have. where, you know, it was like useless metrics and it was like lines of code in a code base, right? number of P PRs, and now it's like the number number of tokens, you know, actually used. These things sound cool though, right? And so that's kind of where it is.

Brian Bell (06:39.472) Right.

Brian Bell (06:57.179) Right. Right.

Jonathan Archer (07:02.742) I don't know. I I do see these numbers of tokens that are being given out in a subscription dropping. So having a gateway where you could actually bundle up multiple subscriptions makes more and more sense. And then what happens?

Brian Bell (07:14.479) And you guys are bundling subscriptions, not just the API, it sounds like. So that's interesting. How does the how does that work? Because usually you have to kind of log in to their portal or or desktop app to kind of use the subscription.

Jonathan Archer (07:17.737) Exactly. So

Jonathan Archer (07:28.108) Yeah, so you still do. and what we do is we actually have a small background process running on your on your machine to make sure that we're fully compliant with the upstream provider's terms of service. So you did briefly mention Gemini in the intro. We don't actually offer any Gemini subscriptions to the routing because that could actually get you banned, okay? And imagine losing your Google account. I mean, I don't know what you have on your Google account, but it could be family photos, it could be, you know, all your drive. You know, losing that account is very scary. So for us, that's a main priority, is like, okay, how do we do everything in the most compliant way with upstream providers? And you'll see you download this background process, and then you do actually have to log into Claude, you have to log into Chat GPT, you have to log into Kimmy.

Brian Bell (07:56.526) yeah.

Brian Bell (08:04.646) Yeah.

Jonathan Archer (08:22.222) Whoever it is, and then we're just checking to make sure that you actually have access to those. But you're never logging in through our portal. You're never giving us your credentials. Because if you were to give us your credentials, it would be a violation of of terms of service, but we're allowed to check that you actually have them.

Brian Bell (08:43.323) Hmm. How do you kind of keep the project memory, you know? I think Cloud's doing a a much better job of this in the last, you couple of months really, where they're starting to store memories and MD f MD files and stuff like that. How do you kind of keep that synced across all those subscriptions?

Jonathan Archer (09:01.378) Yeah, so we have them in our database, and these are all encrypted beh behind a seed phrase with zero knowledge proofs. So this is kind of our crypto background coming into play into this world. but then the memories is like a tricky game. It's like, okay, like how do you choose what becomes a memory? You know, you could have the user explicitly prompt, you know, this should be a memory. but it's like an ever-evolving kind of feature that I would say that. We're closely monitoring the the industry to see what the industry s standards are and like seeing like, okay, like Hermes, for example, does it like this. Okay, we could incorporate some of those features. you know, Claude does it like that, all right, and and kind of mix and match.

Brian Bell (09:46.96) And can you is it in the terms of service, is it okay to kind of because cloud stores all this stuff wherever I tell it to on my my machine, all these memories and stuff. can you just grab those files and kind of, I don't know, look across the different subscriptions and pull that into your open L L layer or?

Jonathan Archer (10:05.709) Yeah, if they're in your machine, you know, totally. And depending on what what you give access Claude to do, right? So sometimes people only give access to a certain folder or file on their machine, and then other times they give access to the full machine. and so I'm also investigating these sort of things right now with my co-founder so that when new users come into OpenLM, how do they just port over all those memories from Claude, from Codex, you know, from Hermes? straight into OpenLM so that whenever they're using Open LM everywhere, it has that unified kind of memory set.

Brian Bell (10:42.673) It's interesting. It kind of reminds me of like Kubernetes, but for LLMs, right? Kubernetes, of course, created these containers that sits on top of, you know, sharded out infrastructure. And so you can basically, you know, drop and run your applications anywhere, right? Is that kind of how you think about it or is it something different?

Jonathan Archer (11:04.471) No, and I love that analogy. I just wrote it down because I think it's really interesting. That's one of the main things that we're actually going towards is this idea of like also multi-device management. So, you know, you know how we just talked about how you do actually have to log in, right? So I have to go into ChatGPT, login, and you know, cloud and et cetera, et cetera. Well, what's really cool about OpenLLM is because we're pairing up these different devices. You could have up to five devices in the Pro Plan. So I just spun up a new Hetzner virtual private server the other day, and I've got another one already running. This one that's already running already has access to all my subscriptions. It's got my Hermes agent running on there. and because I was too lazy to just go and like log in to the new ones on the new machine, it's actually just pulling inference from that other machine. Okay? So it's just checking, hey, This one has the access to the subscriptions. I've got a a daemon, I've got the background process. I could actually just send my request to this machine instead of having to authenticate here. Now, why is that actually really, really bullish? It's because you're gonna see in the future agents spinning up their own machines to spawn subagents just to like get certain jobs done. You know, so when they do that, where are they gonna get their inference from? Well, either API keys or from that subscription from that other machine where they could just get the inference like that. And it's also good not just for cloud providers when getting inference, but also if you want to run your own local model. So if you have your own local model, but you're obviously not gonna run it on your laptop because you don't have the proper hardware to do so, you have it running on this VPS and then you pull inference from there. So it's really, really interesting. And I think it's a great analogy.

Brian Bell (12:58.691) Yeah, and I'm guessing you guys can integrate I think what you're describing also is another version of hosted models too, where I could self host a model and you guys could dynamically route to that, or I could stand up on my own open source on some other infrastructure provider and you would dynamically route it, or is that possible?

Jonathan Archer (13:19.533) Yeah, I mean the the possibilities are limitless, you know, and even down the line, like host your own website, you know, if it's not like a a crazy website with a lot of traction, then you know, you could spin up a server and on a roadmap, that's one of our ideas, is actually to have one click deployment of a VPS to make it really easy for these like less technical users to actually have access to these things. Because like obviously a lot of people are talking about agents and how They have them running 247, but not a lot of people really know how to do that. Or when they go into AWS, it's like, man, this is this is scary. What is an instance? What is EC2? You know, all those sort of things. So we're just gonna have one click deployment and they have like a simple script that runs automatically where it downloads Hermes or OpenClaw or whatever it is, where they could plug in their their subscriptions or get inference from another machine as well.

Brian Bell (14:16.987) Val the value prop here is like, okay, you are paying, you know, two, three hundred bucks a month to one provider. So you can kind of token max and maybe more, maybe thousands if you're a developer. What if you had three twenty dollar subscriptions that could dynamically route between the top three or four providers? So, you know, sixty to eighty bucks, and then you wouldn't have to pay a thousand bucks a month. You could pay open LLM, what however you guys price it. Yeah. So it's thirty bucks a month.

Jonathan Archer (14:43.833) Thirty bucks. Yeah. Thirty bucks a month. yeah, that's it. There's also a free tier, so people can hop in right away and get fifty million tokens, you know, routed just trying it out, up to two devices. But yeah, the it's definitely a possibility. And I think beyond like the tokens, because you know, if you're bundling, let's say five twenty dollar subscriptions.

Brian Bell (14:48.507) To start.

Brian Bell (14:53.977) Kind of try it out.

Jonathan Archer (15:11.491) You're still gonna get way less use usage than having a $200 plan with one provider. Although what I would say that's really beneficial is that each model has their own little features that say that make them really interesting to use. so for example, Grok right now is really great at searching Twitter. You know, and that's like a feature that Grok obviously has that you know these other ones aren't.

Brian Bell (15:35.513) Obviously, yeah. As first party access, so yep.

Jonathan Archer (15:39.918) Yeah, you know, it makes a lot of sense. So when you're developing a marketing plan, okay, it calls Grok to look up Twitter, you know, see what's going on over there. And then you actually call your ChatGPT subscription to generate some images, which it's also great at, you know, and then you've got Claude to actually do the copy because it it sounds, you know, I would say generally speaking, the best.

Brian Bell (15:54.053) Yeah, 'cause they're they're really good at images, yeah. Yep.

Brian Bell (16:02.705) Best best writing, yeah, prose is definitely Claude. That's interesting. So you guys are already doing a little bit of that in the background. So if I have those three or four subscriptions and I do a prompt and open LLM, it's already dynamically routing some of the jobs to the different subscription providers. Wow.

Jonathan Archer (16:19.994) Yeah, definitely. I mean that's something that's really novel about OpenLLM that you don't see is that you actually get to unlock these features in any of the CLIs that you choose to use. So for example, if you're using Claude Code, like I could be generating a landing page and actually use my ChatGPT subscription to generate an image or video, whatever it is, for my landing page directly in that CLI. without ever having to leave or or switch over, you know, which saves you a lot of time and is quite quite novel. We haven't seen this in other places.

Brian Bell (16:56.113) What's been the hardest part product wise or technical wise so far of of building this?

Jonathan Archer (17:03.588) Yeah, definitely this background process, the daemon that I had mentioned earlier, because that had many, many implications. And really, you know, when you're building a product like this and making sure that it's compliant with upstream providers' terms of service, you know, you want to do it also in the most efficient way. And there's also like a bunch of different ways of doing it. There's these hand roll processes or bridging. we do actually a mix depending on the situation. And depending on the provider's terms of service. So for example, like Claude Code allows you know X number of things where Codec is a lot more liberal, right? And so right now you could actually use OpenLLM within Hermes, but once it's once Claude Code sees any Hermes in the header tag, it'll actually re reject those requests. So it's so just making sure we know how to manage and and do all these things in the most compliant way. Yeah. Yeah.

Brian Bell (18:02.779) Claude would be like, I'm sorry, I cannot work with Hermes or something. It's how it's how it's howling us. from can't sorry, Dave, I can't do that. Open the door, Hal. Open the door, Hal. Sorry, Dave, I can't do that. That's funny. yeah, I I see Claude is a little bit more preachy. Do you do you kind of agree as you kind of use these? I kinda I kinda use all of them as well. And I feel like Claude is best writer, but it also gets a little bit a little bit preachy.

Jonathan Archer (18:09.424) it'll just it'll just give you the the the cold shoulder.

Brian Bell (18:32.774) Like

Jonathan Archer (18:33.594) Yeah, sometimes you're like, Yeah, just get to the point, you know, like and and also it's like a a little bit more

Brian Bell (18:36.581) Yeah. It it completely gets this idea that's wrong and then like gets stuck in this idea. And I'm like, no, that's not what I asked you to do. And that's actually like orthogonal to to the conversation. It doesn't yeah, and it's just kind of stuck on this like little thing. and it'll kind of anchor. I find like Claude is a lot more prone to like anchor bias. Like I can actually the way I prompt it, it could lead it. I could I could kind of guide it towards the decision I want.

Jonathan Archer (18:49.264) Sorry.

Jonathan Archer (19:05.668)

Brian Bell (19:06.565) You know, pretty easily by just kind of like anchoring.

Jonathan Archer (19:10.884) Yeah, no, it's it's funny. You know, it's really good to have different models for different use cases. And that's I think one of the best selling points with OpenLM is like, okay, you bring your own keys, you bring your own subscriptions, but we're gonna live in a multi-model future. And I think this is also kind of from my crypto experience, is like, okay, there's gonna be more than one ecosystem, and it's actually the best when you start using different products together, you know, because they have different use cases. And so that's something that we're working on for the landing page is these kind of like cookbooks and recipes, let's say, for you know, what model to use when, because yeah, it makes a lot of sense and it's really annoying if you have to switch, you know, different UIs all the time to get access to each of those models. Like you you don't wanna have to go to Grokbot and then back to I don't know, Claude Code and then back to Codex, you know, just Do whatever you want to do in the place that you get the most work done.

Brian Bell (20:12.197) Yeah. What's tell me more about the go to market. So the you have a freemium version. What's the tripwire where like I kind of run into the limits of the freemium version?

Jonathan Archer (20:24.186) So you get to use up to fifty million tokens a month in the free version. and then obviously when you pay, you have unlimited tokens that you could route. The other thing that's really interesting about the paid version is you have access to more devices. So you could set up up to five devices, where in the freemium or free version you only have two.

Brian Bell (20:45.977) five subscriptions so I could si I I could sign up my gr my my groc and Bod

Jonathan Archer (20:50.808) No, in both of them you have unlimited subscriptions and unlimited API keys. so there's no problem in that. It's more like I've got my laptop running, so that's one device, then I've got my Hetzner, that's another device, and I've got my second Hetzner. It's like AWS. It's like a yeah, virtual private server provider. and I've got those two, so right now I'm at three devices. So if I wanted to have all three of these devices running.

Brian Bell (20:54.624) okay.

Brian Bell (21:03.407) Right. What's a Hexner for people who don't know? Okay. Gotcha.

Jonathan Archer (21:20.836) Then I would need to have a paid version.

Brian Bell (21:23.601) Got it. that's really interesting. And then I I like the the stance of like you're building up the dynamic routing between the different providers and the memory. Right. yeah, I think that's what

Jonathan Archer (21:37.4) Yeah. I would just say like I wanna be careful in terms of the dynamic routing because there is a little bit of it going on, but not so much. Like we're not a smart router, for example. So it's not like when you say like, hey, you know, create a landing page, it knows to call this model right away. Like there there is a little bit of that going on and we are

Brian Bell (21:57.18) Can I prompt it that way? It's like, hey, go go to chat GPT and get the image and then Yeah.

Jonathan Archer (22:01.328) For sure, for sure. And and so we're we're actually looking right now at creating these sort of skills to make sure that it's always using the best ones. Because like also sometimes what happens is that, you know, maybe you have an API key. and so let's say you have your ChatGPT API key and you ask it to generate an image. Well, we want to make sure that it's using your ChatGPT subscription to generate that image instead of calling the API key if it's available to you. You know, so so we are working out the kinks in terms of that system. Really what's going on now is like we've got a fallback chain which is configurable by the end user. So I've been l loving Astra, so I've got it on top of my Ultra and it's divided like Ultra, Plus, and Lite, you know, kind of like Opus Sonnet haiku, you know, based on intelligence levels. and basically what happens if Astra runs out of

Brian Bell (22:52.496) Right.

Jonathan Archer (22:59.786) Usage, then I start using the next in line in my ultra tier, which right now is Fable five point one. Okay. but also it's not even just about usage because sometimes these providers go down. You know, every once in a while they go down for a couple hours. So it's making it called chat GPT, it's down, okay, it starts using Fable and then it keeps going like that.

Brian Bell (23:21.263) Interesting. So you kind of tiered it on the front end where I can like, Hey, this is an ultra intelligence kind of task. So use my ultra tiers across my my subs. this is like kind of important. So use the Opus kind of five point six kind of model and then then so on.

Jonathan Archer (23:31.087) Yeah.

Jonathan Archer (23:38.551) Exactly. And actually we also set it up so you know when it spawns sub agents, it'll start spawning the plus plan, you know, because you don't actually want Fable running around doing these small tasks. You want, you know, to actually use Sonnet or or Haiku in terms of the Claude family. because yeah, otherwise you're basically burning tokens. Like would you would you go send Steve Jobs? You know, to go get a file? No. Yeah, yeah, yeah, yeah. It d it doesn't really make sense, right? So you wanna make sure you have the right person for the right job.

Brian Bell (24:09.753) Steve Jobs open a coffee shop.

Brian Bell (24:19.185) That's awesome. So I mean, this there's a lot of infrastructure in AI right now. There's an open source gateway that does something similar with tens of thousands of GitHub stars. I'm sure you guys thought about going open source, but how do you kind of stand out in a crowded field?

Jonathan Archer (24:35.652) Yeah, so we're actually we're we're in an interesting situation because we're half open source. So we we've got half maybe is a stretch, but we have some of our product which is open source because we have this background process running on your computer. and I told my co-founder, hey, let's make this open so that people could verify that this is like completely legit and they could contribute if they want to. What's closed is actually our front end and dashboard. and this unified dashboard also makes the setup so much easier. You know, it allows people to see what's going on. They've got a chat window, adds a bunch of extra features. And then like how do we stand out? I think it's just like really being aggressive in terms of listening to product features and requests from the community to to make them like be seen and heard. You know, and and that's what we're doing. We're talking to everyone as much as possible. I even would like to set up a like a friendly AI who could chat and hop on calls, right? And just like take product requests, you know, have an agent start working on the code and then we review that PR and and push it if need be. My my co founder might might kill me for saying something like that because he always gets worried when agents get in the loop, you know. But yeah, we are becoming more and more agent first.

Brian Bell (26:07.003) What are the top three features on the roadmap for the next, you know, six to twelve months?

Jonathan Archer (26:12.912) So we're really focused on multi-device sessions and what that looks like. I think that's gonna be very, very interesting. and so right now we even have one feature that's that's basically ready. We just want to make sure we've got enough testing and security checks going. And that's being able to use the UI to tunnel into a computer and then actually access cloud code or Hermes or whatever it is to actually start writing code from the UI. So I mean let's say you're on a train, you know, in in Europe and whatever you don't want to pull out your laptop. So you just pull out your phone and you could actually start shipping code like that. right now, because of the shared inference, you could use your phone and have like an integrated chat feature, which is really cool where you could actually chat with all the different models. But yeah, the the multi-device is the future. and then beyond that, we're also looking at how do we how do we best serve enterprise? enterprise is obviously a huge market. They can't use subscriptions, right? So when you're an enterprise level, you actually do have to use API keys. But we do have amazing, amazing routing with API keys and we don't upsell tokens like other providers do. So yeah, we're we're we're looking into that and how how the multiplayer AI world is gonna play out, right? Because having a shared memory and indexed code base between members of an organization I think is very interesting and and seeing how they could work together with multi devices is cool.

Jonathan Archer (28:03.28) Sorry, you're just on mute.

Brian Bell (28:06.681) My dog was barking. what what's an idea in AI right now that you think is wrong?

Jonathan Archer (28:15.076) Yeah, this is an interesting one. I think something that I'm I'm seeing is a lot of people trying to restrict people to their own harness, you know, and like lock them in. And so with Open LLM, that's like the opposite of what we're trying to do. Yeah, you could get a lot done within Open LM, within these chat windows, you know, as I mentioned, future plans, but we also want people to be able to use their tokens wherever they want. You know, it's their tokens, they should be able to go. So we've seen some platforms and harnesses that also do some sort of subscription routing, let's say, but it's only within their closed system, you know, which I think is a drag because there's a lot of new tools that are coming out that are amazing. And people also love the tools that they're already using. For example, if you're a codex guy, like just keep using codex, you know, I don't I don't think it should stop you. but you're gonna need to have like access to as many tokens as possible. And that's what open LM is is allowing you to do. It's okay. Have access to all your tokens, all the newest models, try it try them out, but use them in the environment that you already feel comfortable using, right? And that that could be codex, could be cloud code, whatever it is.

Brian Bell (29:36.569) What's a gateway layer that you look up to in the past? Is there some startup or company where you're like, Man, if we could just do what that gateway layer did, you know, five, ten, twenty years ago, that'd be amazing.

Jonathan Archer (29:53.381) Yeah, I mean open router is a huge inspiration in terms of what they've done. Also

Brian Bell (29:59.952) Expla yeah, explain what they did to the audience.

Jonathan Archer (30:03.192) So open router, I would say, was the first big w gateway, you know, and they they've just been acquired by Stripe, you know, for $7.5 billion, which is quite incredible considering it's only a three, maybe four year old company. I'm not sure. and you know, I think their their their go to market was really interesting in terms of just showing people which models are being requested, you know. giving access to so many different models. They're also from the crypto background. I think the CEO was the previously the CTO of OpenC, which was like one of the largest NFT marketplaces at one point. And yeah, just like that outlook, you know, really made a lot of sense to us. the thing is is that they've created this whole business model around API keys and it's great for enterprise and you know, vibe coders But they couldn't really ever access this subscription route, which we're we're focused on, because it would kind of tarnish the relationships, I believe, with these upstream providers. So they have they have what I I understand to be like two parts of their business is that they're buying tokens in bulk. So they're probably getting a discount on the tokens from the Claude, Chat GPT, whoever. And then they also sell f they they put a five percent, let's say, tax. on whatever it is you're buying. So if you buy $100 worth of tokens from Claude using Open Router, you actually have to pay $105, right? Whereas, you know, with with Open LLM, it's bring your own keys. So if you have an API subsc or key from Claude, pay $100 to them. Well it just costs you $100 when you use it through Open LLM, other than our subscription.

Brian Bell (32:01.041) Why do you think that was worth seven billion dollars to Stripe? That's kinda crazy.

Jonathan Archer (32:06.816) Yeah. I I was I was actually just looking into that a little bit more last night and I would like to say I I did a bit of a deeper research, but I think it was it was really they they have a very, very large user base, you know. they're the number one reseller of tokens in the world. they've done a ton, a ton of volume and there there's gotta be a lot of interesting data, you know, that they've got Which which I imagine is the real reason why they're getting bought at such a high price. Kinda like Cursor. Cursor got bought not for not for their VS code fork, but for all the data that they have.

Brian Bell (32:50.425) And so do you see what do you think is the likelihood of open L LLM getting acquired by one of the big hyperscalers? Because, you know, eventually this gateway layer probably gets you know gets compressed down to different platforms, right? With Stripes trying to do it because they kind of see themselves as a platform, a good like a gateway layer on payments. So they're like, okay, we'll become a gateway layer on AI. I think All the hyperscalers would probably want to do the same thing, especially, you know, I I I worked at Azure and I worked at AWS and I can imagine them looking at open router and they probably did have a bid on that acquisition. Cause, you know, their their whole thing in their ecosystems and Google Cloud to a certain extent as well, is like how do we drive consume revenue underneath? And, you know, so eventually it's gonna be like, yeah, you you know, we just have this like layer, That sits on top of the cloud and we can do that for you. But all these models run on our cloud. Any model you want. You want to run cloud, you know, and that's on bedrock. You know, you want to run chat GBT, op op open AI on bedrock. how do you see this kind of industry kind of unfolding over the next few years?

Jonathan Archer (34:09.156) Yeah, I think there's I mean, to answer the first question, you know, I think there's a very high chance that OpenLM one day could get acquired by a by a big player as well. And I believe there's a lot of alignment between us and like Vercell. That's kind of the vision, and that's because we were built largely using Vercel, and we've seen even they are also doing a kind of gateway business where they aren't upselling, but they're just trying to offer the best. possible service to their consumers, right? But again, they're doing it through API keys, right? And yeah.

Brian Bell (34:43.505) Dorsel is basically a gateway, right? They're probably built, I don't know for sure, but I'm guessing they're built on one of the hyperscalers. And they kind of disintermediate them and and make it really simple and easy to stand up infrastructure and monitor it and do all your DevOps and and stuff like that. I mean, we we run Team Ignite on Drusel, so pretty familiar with it. But I could I could easily go to AWS, but it's a lot harder to use and clunky and very sophisticated. Yeah, it's just

Jonathan Archer (35:00.868) Yeah.

Jonathan Archer (35:07.368) it's such a drag. It's it feels like such an outdated system, you know, it's like I it's like poking around on on Oracle, you know, it's like what what am I doing here?

Brian Bell (35:19.289) trying to run yeah trying to run your your CRM on like Microsoft dynamics. It's like yeah it's like really powerful but really clunky to use.

Jonathan Archer (35:22.264) I yeah, you know, like the new the new age, they want like it to be accessible via CLI, right? Like I wanna be able to go in cloud code. Like I love Stripe actually because like I could do everything from Cloud Code and Stripe, basically. You know? Like I'm talking to you and I'm like, hey, we should we should create a coupon code for the Ignite community. Well I just go on c on cloud code, I'm like all right, ignite eighty. You know, and give eighty percent off for two months to the Ignite community, which we could actually do. I think that'd be fun. but yeah, this these sort of things are are amazing. And when you enter these like older systems, it's just like it's overly complicated. There's like a million tabs, you just get lost. It's quite daunting. And so this is even the yeah.

Brian Bell (36:12.881) Yeah. They got like a thousand services. You're like I don't even know where to to get started here. Yeah.

Jonathan Archer (36:18.532) And this is an ethos that we've actually adopted within Open LLM. And I was just talking to my co-founder about it before this call. It's like, yeah, I love how with Open LLM, one, a new model comes out. We have something called dynamic discovery. So we actually are querying these model providers all the time to see what models are available. And so sometimes like we actually see the model on Open LLM before they tweet about it because we're querying it all the time, which is pretty cool. and so let's say you know this new model comes out, Astra six. Instead of me having to go to my OpenLLM UI and putting it at the top of my fallback chain of Ultra, I just go into Cloud Code and I say, hey, can you can you pop in Astra six at the top of my Ultra plan and and it'll do that for me? You know, or or hey, what's the usage? You know, how much how much time left? how many tokens do I have left on my ChatGPT plan? You know, these sort of things. And that's really the future and you know, Vercell's been so so good at that, as well. So yeah, I believe there's a lot of synergy between us and them.

Brian Bell (37:28.219) So five years out, do you think the gateway layer is still a standalone product? Or does it get absorbed into the clouds or model providers or agent frameworks, et cetera?

Jonathan Archer (37:38.991) No, it can be standalone. Even today, like that's why we're not we're not focused on just being a gateway. We actually want to be all your AI infra in one place. so we started out as a gateway, is what I like to say. And and we're going towards this all your AI infra in one place. Because yeah, it's just not enough. You know. That's why like I'm so excited about the multi-device management. tunneling in between devices, sharing inference between devices, you know, these sort of things are are really what's interesting. And I believe where the industry has had it

Brian Bell (38:19.313) Yeah, awesome. Well let's wrap up with some rapid fire questions. what's the best advice you ever received?

Jonathan Archer (38:29.872) Don't do drugs. Yeah.

Brian Bell (38:36.101) What about advice you initially ignored but turned out to be really good advice?

Jonathan Archer (38:41.008) Don't do drugs. No. No. yeah. Also I'm kinda I'm kinda drawing a blank, but we we could keep those in, it's funny.

Brian Bell (38:51.109) Yeah, yeah. What's what's something you believed with total conviction five years ago that you'd argue against now?

Jonathan Archer (39:00.506) Yeah, I mean coming from that NFT craze, I would say like, you know, at one point I never thought it was gonna be like hanging up pictures of monkeys on your wall and like that was the be all end all. But I I did believe that, all right, you know, this will be passports, IDs, you know, deeds, etc. And I still do believe there's great use cases in those, but it's just hard for adoption, right? Sometimes the tech is there, but the adoption just isn't. so yeah, that would be it.

Brian Bell (39:31.58) Problem with crypto and and I've inv invested in some crypto stuff over the years. Out of 400 portfolio companies, I probably have a handful that are actual crypto related, but they're they tend to be picks and shovels kind of infrastructure stuff. Like, okay, well, how do we, you know, make them more secure? Or how do we help governments manage the seized assets of crypto? my problem with crypto always was like, it was always like, yeah, it's better.

Jonathan Archer (39:40.656) Mm-hmm.

Brian Bell (40:00.346) and I'm like, okay, great. Then go implement it. And it was like a hammer looking for a nail, right? It's like, well, why why do we need that that way? And I feel like like if we had the if we were building out all the financial systems now, like if we're rebuilding them from scratch, on a new planet or something, we'd of course we'd use like blockchain and stuff. But it's so hard to rip people out of what's already working.

Jonathan Archer (40:10.553) Right.

Brian Bell (40:28.069) And all the network effects of all the payment systems and banking systems and stuff like that. I think I think what's got a foothold recently, and you know this space a lot better than I do. so you can weigh in here, but is the stable coins, right? That's been kind of the thing that's really taken off. You know, besides like Bitcoin for a store of value, some smart contract stuff, but really the stable coins, right?

Jonathan Archer (40:50.18) Yeah, and I mean not to get like too deep into like all these conspiracy and whatever sort of things, but you know, it's really not within these greater powers' interest, you know, to push that narrative, you know. And like now they are like whatever the big one. It's like it's like these cl yeah, exactly. You know, it's like these closed systems. And that's why you're also seeing across a lot of nations that they're even like trying to remove cash because they don't want people to be able to

Brian Bell (41:04.831) yeah, like that's their that's their bread and butter. That's their cash cow. They own the payment rails, right? Yeah.

Jonathan Archer (41:20.08) pay each other in like a in like a yeah or or just in like a peer-to-peer way. Like everyone wants their piece of the of the pie, right? And so for me to send you money, you know, is quite complicated beyond like I don't know, use Western Union and takes time. Or I don't know, I use like Stripe. Like I I maybe you have Revolute and I don't, you know.

Brian Bell (41:23.131) Untraceable currency. Yeah.

Jonathan Archer (41:46.019) I I have to use wise and it's a bank wire. Okay, send you a hundred dollars cost me like five bucks. I like these sort of things. It's like ridiculous. But so many people are getting paid along the way. Whereas like when I'm using crypto, like I send you money, you get it. You know, there's no one approving it or denying it. and it can be done in like a traced way. Like it doesn't have to be like, it's all private, like I'm just trying to like launder money. No, like it I could Show the government, hey, this is my money. I'm sending this guy money. It's all good. But it's just the fastest and cheapest way to do it. You know, which is like what I love. And you also see it. I've been living in South America for a long time, and I I lived in Argentina for quite a bit. So that's why I'm drinking some mate every once in a while. and the the crypto adoption in Argentina is so, so high because. Of inflation, right? And because people just needed some sort of alternative. And what's really interesting about Argentina is that they have more US dollars per capita than the US, like physical dollars, because of the amount of people who are just like storing physical dollars in fear of like hyperinflation or whatever it is. so storing u physical US dollars is

Brian Bell (42:56.045) wow.

Jonathan Archer (43:09.328) pretty impractical, you know, and quite dangerous. Like what if you have a fire or robbery or whatever? Having having USDT, like you mentioned, is is quite amazing. You know, it's a great way and it's very it's very portable, you know. it's it's really the best Rails. And I believe actually now seeing where we're headed is maybe crypto wasn't really created for us. Maybe it was actually created for the agenda community. you know, for agents to transact. Because for an agent to spin up a wallet and send money, you know, it's it's the best and fastest way. You know, and they do it like that. It's like very easy for them to understand.

Jonathan Archer (43:57.262) I don't know, does that that answers your question, right?

Brian Bell (43:57.657) If everything Yeah, no, that's great. Yeah. If everything disappeared tomorrow, what would you rebuild first?

Jonathan Archer (44:10.36) I mean I would I would just keep doing what I'm doing today. I mean, if it's everything in terms of my work, if it was like all of everything in terms of the world's work, then there's there's other greater projects to work on. but I I've got full conviction in Open L L and so I would probably just start to to try to rebuild it and keep going. I think there's still enough time to to carve out your space within this sector. So

Brian Bell (44:35.739) Yeah. well really enjoyed the conversation. Where can folks find you online?

Jonathan Archer (44:41.796) Yeah, LinkedIn, you just look up Jonathan Archer, OpenLLM. I am starting to put out some content on Instagram, which is completely cringe. I feel awful doing it, but I feel like this is part of being a SaaS founder these days, is you gotta put out this short form content. So I'm trying it. So it's Johnny J O N N Y dot O L L So for open L L When you're actually using Open LLM on your on your computer, that's one shorter way of using it. Instead of writing Open L L you just write O L L And yeah, that's that's it.

Brian Bell (45:22.617) I appreciate it. Thanks so much.

Jonathan Archer (45:24.251) Yeah, I appreciate it. Thanks, Brian.

This article is for general informational purposes only and does not constitute investment, legal, tax, or accounting advice, nor an offer or solicitation to buy or sell any security or investment product. Investing involves substantial risk, including possible loss of principal, and past performance is not indicative of future results. Full disclaimer.

Subscribe to Ignite Insights

Founder and investor interviews from the Ignite Podcast, the Last Week Ignite weekly market digest, and original essays on venture math, AI, fundraising, and go-to-market — from a seed fund making more than a hundred investments a year.