Thank you. Hello, and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. 90% of the information transmitted to human brains is visual.
So while advances related to large language models and other language processing technology have pushed the frontier of AI forward in a hurry over the past few years, visual information is integral for AI to act with the physical world. which is where a computer vision comes in.
RoboFlow empowers developers of all skill sets and experience levels to build their own computer vision applications.
The company's platform addresses the universal pain points developers face when building CV models, data management to deployment.
RoboFlow is currently used by over 16,000 organizations and half the Fortune 100, totaling over one million developers.
They're a member of NVIDIA's inception program for startups.
RoboFlow co-founder and CEO Joseph Nelson is with us today to talk about his company's mission to transform industries by democratizing computer vision.
So let's jump into it. Joseph, welcome, and thank you for joining the NVIDIA AI podcast.
Thanks so much for having me. I'm excited to talk CV.
There's nothing against language, love language, but there's been a lot of language stuff lately, which is great.
But I'm excited to hear about RoboFlow. So let's jump into it.
Maybe you can start by talking a little bit more about your mission and what democratizing computer vision means and making the world programmable.
At the highest level, as you just described, the vast majority of information that humans process happens to be visual information.
In fact, I mean, humans, we had our sense of sight before we even created language.
It's how we understand the world. It's how we understand the things around us.
It's how we engage with the world. And because of that, I think there's this massive untapped potential to have technology and systems have visual understanding to a similar way that humans do all across really we say we say the universe so when we say our north star is to make the world programmable what we really mean is that like any scene and any thing will have software that understands it.
And when you have software that understands something, you can improve that system.
You can make it be more efficient. You can make it be more entertaining.
You can make it be more engaging. I mean, at RoboFlow, that's like, we've seen folks build things from understanding cell populations under a microscope, all the way to discovering new galaxies through a telescope.
And everything in between is where vision and video and understanding comes into play.
So if this AI revolution is to reach its full potential, it needs to make contact with the real world.
And it turns out the real world is one that is very visually rich and needs to be understood.
We build the tools, the platform, and the community to accelerate that transition.
Maybe you can tell us a little bit about the platform, And kind of within that, how your mission in Northstar kind of shaped the way you develop products and build out user experiences.
And I should mention, Great shout out from Jensen in the CES keynote earlier this year for you guys.
And you raised Series B late last year. So I want to congratulate you on that as well before I forget.
I appreciate it. Yeah, it's a good start.
As you said, a million developers, but there's 70 million developers out there.
There's billions that will benefit from having visual understanding.
In fact, in that Jensen shout out, just like maybe the sentence or two before he describes some of the visual partners that are fortunate to work with the NVIDIA team, He described that the way Nvidia sees it is that global GDP is $100 trillion, and he describes that visual understanding is like $50 trillion of that opportunity.
So basically half of all global GDP is predicated on these operationally intensive visual visual centric autonomy based sorts of use cases.
And so the level of impact that visual understanding will have in the world is just a fraction of what it will look like as we progress.
Now, in terms of how we think about doing that, so it's really about empowering the builders and giving the certainty and capability to the enterprise.
So for example, anyone that's building a system for visual understanding often needs to have some form of visual input, camera, video, something like this.
Sure. You need to have a model because the model is going to help you act, understand and react to whatever maybe actual insight you want to understand.
And then you want to run that model somewhere.
You want to deploy it. And commonly, you even want to chain together models or have a system that triggers some alerts or some results based on information that it's understanding.
So RoboFlow provides the building blocks, the platform, and the solutions so that over a million developers and half the Fortune 100 have what they need to deploy these tools to production.
And you're doing it, kind of mentioned in the intro, We're trying to make the platform available for folks who are deep into this, have been doing CV and working with machine learning for a while, and then also folks who might be new to this, they can get up and running and work with CV, build that into their toolkit.
Yeah, the emphasis has always been kind of on someone that wants to be a builder, that definition is expanding with the capabilities of code generation, prompting to apps.
We've always kind of been bent on this idea that those that want to create, use, and distribute software.
What's funny is that when we very first launched some of our first products, ML teams initially were kind of like, oh, I don't know, this seems pretty pedestrian, I know exactly what to do.
And fast forward now, and it's like, whoa, a platform that's fully featured that has immediate access to the latest models to use on my data in the contexts of where I couldn't even anticipate.
So it's been kind of, as the platform has become more feature rich, we've been able to certainly enable a broader swath of both maybe domain experts of a capability.
But I think broadly speaking, the places that we see the rapid most impactful adoption in some ways is actually bringing vision to others that otherwise may not have had it.
Like what used to be maybe like a multi-quarter PhD thesis level investment now can be something that a team spins up in an afternoon.
And that really has this kind of demand begets demand paradigm.
I mean, for example, like one of our customers, they produce electric vehicles.
And when you produce an EV, you know, inside their general assembly facility, There's all sorts of things that you need to make sure you do correctly as you produce that vehicle.
From the worker safety who are doing the work to the machines that are outputting Say when you do what's called stamping, where you take a piece of steel or aluminum and you press it into the shape of the outline of the vehicle, you could have potential tears or fissures or Remember when you actually assembled the batteries out of the correct number of screws?
Basically, every part of building a car is about visually validating that the thing has been built correctly so that when a customer drives it, they can do so without any cause for pause.
And just a little bit of computer vision goes a really long way in enabling that company and many others to very quickly accelerate their goals.
In fact, this company had the goal of producing 1,000 vehicles three years ago and they barely did it.
They did 1,012 in that year. And then they scaled up to 25,000 and now 50,000.
And a lot of that is on the backs of having things that they know they're building correctly.
And so we're really fortunate to be a part of enabling things like this.
So it's kind of like, you can think about it of a system that doesn't have a sense of visual understanding, Adding images a little bit of visual context totally rewrites the way by which you manufacture a car.
And that same revolution is going to take place for lots of operationally intensive processes, but any kind of place that you interact with the world each day. so kind of along those lines what to use the phrase untapped opportunities you see out there what's what's uh you know the low-hanging fruit maybe the high-hanging fruit but you're just excited about when it comes to developing and deploying computer vision applications.
And we've been talking about it, but obviously talk about RoboFlows. work not just supporting but helping developers kind of unlike helping builders you know unlock what's next The amazing thing is actually the expanse of the creativity of developers and engineers.
It's like if you give someone a new capability, you almost can't anticipate all the ways by which they'll bring that capability to bear.
So for example, I mean, we have folks that hobbyists that'll make things like the number of folks that make things that like measure the size of fish because I think they're trying to prove to their friends that they caught the biggest fish.
And then you separately have government agencies that have wanted to validate the size of salmon during migration patterns.
This primitive understanding size of fish both has what seems to be fun and very serious implications.
Or folks that, I don't know, like a friend of mine recently was like, hey, I wonder how many cars out in San Francisco are actually Waymo's. versus like other sorts of cars.
And what does that look like? What does that track over time?
And so they had a pretty simple like Raspberry Pi camera, parked it on their windowsill and In an afternoon, now they have a thing that's counting, tracking and keeping tabulation on how many self-driving cars are making their way on the road, at least sampled in front of their house each day.
Right. No, I don't want to call anybody out, but that's not the same person who had the video of all the Waymos in the parking lot in the middle of the night in San Francisco going in circles.
No. Okay. Yeah. It wasn't that guy. Yeah, but I mean, the use cases are expansive because...
I don't know, the way we got into this is we were making AR apps actually, and computer vision was critical to the augmented reality understanding the system around us.
And we've since had folks that make board game understanding type technology, D&D dice counters, telling you the best first movie you should play in Catan.
And so basically like you have this like creative population of folks or this one guy during the pandemic, you know, he's really bored, locked inside.
And he, uh, he thought maybe his cat needed to get some more exercise and, He created this system that like attached a laser pointer to a robotic arm.
And with a little bit of vision, he made it so the robotic arm consistently points the laser pointer 10 feet away from the camera.
The cat's like jumping around the living room and makes this whole YouTube tutorial.
But then like the thing that's really interesting, right, is that like You know a technology has arrived when a hacker can just build something in a single setting or maybe in a weekend. you know that what used to be this far difficult to access capability is now broadly accessible.
And that fuels like a lot of like the similar sort of enterprise use cases.
Like we kind of have a joke at Rebel Flow that like one person's hobbyist projects is another person's entire business.
Yeah. So the low-hanging fruit, frankly, is everywhere around us.
Any sort of visual feed is untapped. The sort of images... that someone might be collection and gathering.
I mean, the similar things that give rise to machine learning certainly apply to vision where the amount of visual inputs doubling year on year in petabytes of visual information to be extracted.
So it's kind of like, if you think about it, you can do it.
That makes me think of an episode we did recently with a surgeon who founded a surgical data collective. and they were using um just these i i say stacks they weren't actually videotapes i'm sure but all of this unwatched unused footage from surgeries to train a model, to help train surgeons how to do their jobs better.
But it made me want to ask, Are the visual inputs that can go into RoboFlow, it doesn't have to be a live camera stream.
You can also use archive footage and images.
That's correct. Someone may have a backlog of a bunch of videos.
For example, we actually had a professional baseball team where they had a history of a bunch of their videos of pitching sessions. and they want to run through and do a key point model to identify various poses and time of release of the pitch and how does that impact someone's performance over time.
And so they had all these videos that from the past that they wanted to run through.
And then pretty soon they started to do this for like minor leagues where you might not have scouts omnipresent and you certainly don't have broadcast.
So you just have like this like kind of low quality footage on like cameras from various places and being able to produce like sports analytics out of this information that's just locked up otherwise and this unstructured visual capture is now available for, in this case, building a better baseball team.
Yeah, that's amazing. Building community is something that is both vital to a lot of companies, a lot of tech companies and developer platforms and such.
But it also can be a really hard thing to do, especially to build, you know, an organic, genuine, robust community.
Serving enterprise clients, let alone across this seemingly endless sort of swath of industries and use cases and such, you know, also pretty resource intensive.
How do you balance both? How is RoboFlow approaching building that community that you're talking about just now with serving you know, these, I'm sure demanding in a good way, but, you know, demanding enterprise clients.
I think the two actually go hand in hand more than many would anticipate.
When you build a community and you build a large community, set of people that are interested in creating and using a given platform, you actually give a company leverage.
Basically, the number of people that are building, creating, and sharing examples with RoboFlow from a very early day made us seem probably much bigger than maybe we were or are.
And so that gives a lot of trust to enterprises.
You want to use something that has gone through its paces, been battle tested, something that might be like an industry standard.
And you don't become an industry standard by only limiting your technology to a very small swath of people. enable anyone to kind of build, learn the paradigm and create.
Now, you're right that both take a different type of thoughtfulness to being able to execute on.
So in the context of community building and making products for developers, A lot of that I think stems from, you know, as an engineer, there's products that I like using and the ways that I like to use those products.
And I want to enable others to be able to have a similar experience with the products that we make.
So it's things like providing value before asking for value. having a very generous free chair, having the ability to highlight the top use cases.
I mean, we have a whole like research plan where if someone's doing stuff on a dot edu domain then they have increased access to gpus robo has actually given away over a million dollars of compute and gpu usage for open source computer vision projects And, you know, we actually have this, it's kind of a funny stat, but 2.1 research papers are published every day citing RoboFlow.
And those are things like people that are doing all sorts of things.
That's a super cool stat, I think. Yeah.
Yeah, I mean it just gives you the context to like, that is someone's maybe six or 12 month thesis that they've spent trying to and to be able to empower folks to realize what's possible. kind of the fulfillment of like our mission at its core.
Like the impact of visual understanding is bigger than any one company.
And anything that we can do to allow the world to see, expose and deploy that is important.
Now, on the enterprise side, what we were really talking about is building a successful kind of go to market motion. and making money to invest further in our mission.
And enterprises, as you alluded, are very resource intensive in terms of being able to service those needs successfully.
Even there though, you actually get leverage by seeing the sorts of problems, seeing the sorts of fundamental building blocks, and then productizing those building blocks.
There have been companies that have come before RoboFlow who have done a great job of being very hands-on with enterprise customers and productizing those capabilities.
Companies like Pivotal or Palantir or These large companies that have gone from, hey, let's kind of do like a bespoke way of making something possible and deploying it more broadly.
Now, we're not fully like for like with those businesses.
I more give that as an example to show As someone that is building tooling and capabilities, worst case is you're giving the enterprise substantially more leverage.
And certainly best case is There's actually a symbiotic relationship between enterprises being able to discover how to use the technology, be able to find guides from the community, be able to find models they want to start from.
I mean, Robofill Universe, which is the open source collection of data sets and models, is the largest collection of computer vision projects on the web.
There's about 500 million user labeled and shared images and over 200,000 pre-trained models.
That's used for the community just as much as enterprise.
When you say enterprise, enterprise is people.
There's people inside those companies that are creating and building some of those capabilities.
Now, operationalizing and ensuring that we deliver the service quality, it's just the types of teams you build and the way that you prioritize companies to be successful.
We're really fortunate that we're not writing the playbook here.
There's been a lot of companies that Mongo or Elastic or Twilio or lots of post IPO businesses that have shown the pathway to both building really high quality products that developers and builders love to use and ensuring that they're enterprise ready and meeting the needs of high scale, high complexity, high value use cases.
So you use the word complexities and you know, one of the things that I hear all the time from And I'm sure you more than me from people are trying to build anything is sort of how do you balance, you know, creativity and coming up with ways to solve problems.
And particularly if you get into kind of a unique situation and you need to find a creative answer with things. with not letting things get too complex.
And in something like computer vision, I'm sure the technical complexities can spin up in a hurry.
What's been, you know, your approach and what successes, how have you found success in balancing that?
Complexity for a product like Reliable Flow is always a balance.
You want to offer users the capability in the advanced settings and the ability to make things their own while also part of the core value proposition is simplifying And so you often think, oh man, those two things must be at odds.
How do you simplify something, but also serve complexity?
And in fact, they're not, especially for products like RoboFlow, where it is for builders, you make it very easy to extend.
You make it very interoperable. meaning there's open APIs and open SDKs, where if there's a certain part of the tool chain that you want to integrate with, or there's a certain specific enterprise system where you want to read or write data to,
That's all supported on day one. And so if you try to kind of boil the ocean of being everything in the platform all at once on day one, then you can find yourself in a spot where you may not be able to service the needs of your customers Well, in fact, it's a bit more step-by-step, and that's where the devil's in the details of execution of which steps you pick first, which sort of problems you best nail for your customers,
But philosophically, it's really important to us that, for example, when someone is building a workflow, which is the combination of a model and some inputs and outputs, You might ingest maybe like an RTSP stream from like a live feed of a camera.
Then you might have like a first model that's, Let's say the problem that we're solving is we're an inventory company and we're concerned about worker safety.
You might have a first model that's just like constantly watching all frames to see if there's a presence of a person, a very lightweight model kind of running the edge.
And then maybe a second model of when there's a person, you ask a large vision language model, a large VLM, hey, is there any risks here?
Is there anything to consider? Should we look more closely at this?
And then after the VLM, you might have another specific model that's going to do validation of the type of danger that is interesting or maybe the specific area within your your store, maybe you're going to connect to another database that exists.
And then like based on that, you're going to write some results somewhere.
And maybe you're going to write that result to have insights of how frequently there was a cause for concern within the process that you're monitoring. just as much as maybe you're flagging an alert and maybe sending a text or an email or writing to an enterprise system like SAP to keep track.
And at each step of that juncture, any one of those nodes, since it's built for us on an open source platform, which we call Inference, you can actually mix and match, write your own custom Python, write an API in one way or another.
Let's imagine a future where someone wanted the ability to write to a system that we didn't support yet. like first party, you're actually not out of luck.
As long as that system accepts a post request, you're fine.
And so you have the ability to extend the system.
And so it's this sort of paradigm of interoperability and making it easy to use alongside other tools.
And it gets back to your point around servicing builders just as much as the enterprise, I actually think those things are really closely interlinked because you provide the flexibility and choice an ability to make something mine and build a competency inside the company of what it is I wanted to create and deploy.
The way you frame that makes a lot of sense and makes that link very clear.
We're speaking with Joseph Nelson. Joseph is the co-founder and CEO of RoboFlow.
And as he's been detailing, RoboFlow provides a platform for builders to use computer vision in what they're building.
Joseph, you know, in the the intro alluded to all the advances and buzz around large language models and that kind of thing over the past couple of years.
And I meant to ask, Roboflow was founded in 2020?
Roboflow Inc. was incorporated in 2020. That's right.
Got it. And so anyway, kind of fast forwarding to more recently, the past, I don't know, six months, year, whatever it's been.
A lot of buzz around agents, the idea of agentic AI, and then, you know, There was buzz, I guess that the word multimodal was being flung around, uh, kind of more frequently, at least in circles I run in for a while.
And then it. sort of dropped off just as people, you know, there were the consumer models that the clods and chat GPTs and Geminis and what have you of the world.
Just start incorporating visual capabilities both to, you know, ingest and understand. and then to create visual output, voice models, now getting into short video clips, all that kind of stuff.
What's your take on the role of multimodal AI integration when it comes to advancing CV how is RoboFlow positioned to support this?
Multimodality allows an AI system to have even more context than from a single modality, right?
So if one of our customers is monitoring an industrial process, And let's say they're looking for potentially a leak, maybe in an oil and gas facility.
That leak could manifest itself as, yes, you see something, a product that's dripping out and you didn't expect it to.
It also could manifest itself as you heard a noise.
Or maybe there's something about the time dimension of the video that you're watching as another modality beyond just the individual images.
And those additional dimensions of data allow the system that you're building to have more intelligence.
And so that's why you see all these modalities crashing together.
And what it does is it enables our customers to have even more context.
The way we've thought about that is, We've actually been built on and using multimodality as early as 2021.
So in 2021, there was a model that came out from OpenAI called CLIP, Contrastive Language Image Per Training. which introduced this idea of training on 400 million image text pairs.
Can we just associate some words of text with some images?
What this really unlocked for our customers was the ability to do semantic search.
Like I could just describe a concept And then I can get back the images from a video frame or from a given image that would be interesting for me for the purposes of building out my model.
Increasingly, we've been excited by increases of models that have more multimodal capabilities on day one.
That comes with its own form of challenges though. the data preparation, the evaluation systems, the incorporation of those systems into the other parts of the pipeline that you're building.
And so where there's opportunity to have even more intelligence, There's also challenge to incorporating that intelligence, adapting it to your context, passing it to other sorts of systems.
And so RoboFlow, and being deep believers in multimodal capabilities very early on have continued to make it so that users can capture, use, and process other modalities of data.
So for example, we support the ability for folks to use vision language models, DLMs, in the context of problems they're working, which is typically like an image text pair.
So if you're using, you know, Quen VL 2.5, which came out last week, or Florence 2 from Microsoft, which came out maybe about six months ago, or PolyGemma 2 from Google.
These are all multimodal models that have very rich text understandings and have visual understandings, which makes them very good at, for example, document understanding.
Like if you just pass a document, there's both. text in the document and a position in the document.
And so RoboFlow is one of the only places, maybe the only place where you can fine tune and adapt, say, QuenVL today. which means preparing the data and running it in the context of the rest of your systems.
Those sorts of capabilities, I think, should only increase and enable our customers to get more context more quickly from the types of problems that they're solving.
So I think a lot of these things kind of like are crashing together into just like AI, like amorphous AI that has all these capabilities like you'd expect it.
Yep. But as that happens, what's important is there's actually still unique parts of visual needs, right?
Like visual needs require visual tooling, in our opinion.
Like you want to see, you want to validate, you need to do, you know, the famous adage of a picture being worth a thousand words is extremely instructive here.
Like you almost can't anticipate all the ways that the world's going to look different than how you drew it up like self-driving cars are kind of this example 101 where yeah you think you can drive like you have a very simple way of describing uh what the world's going to look like but Don't know like let's take a very narrow part of a self-driving car stop signs, right?
Okay, so go stop signs look universal. They're always occupied They're red and they're really well mounted on the right side of streets.
Well, what about a school bus where the stop sign kind of flips off?
Where it comes on. Or what about like a gate of where like the stop signs mounted on a gate and the gate could open and close.
And pretty soon you're like, wait a second.
There's a lot of cases where a stop sign isn't really just a stop sign.
And seeing those cases and triaging and debugging and validating, we think inherently calls for some specific needs for processing the visual information.
And so we're laser focused on enabling our customers to benefit from as many modalities as help them solve their problem, while ensuring the visual dimension in particular is best capitalized on.
Right. And I may be showing the limits of my technical understanding here. have added if so but does that exist as you know robo flow creating these you know sort of as you said, amorphous AI all crashed together models that have this focus and these advanced visual capabilities, or is it more of a chaining a RoboFlow specific model onto other models?
Commonly, you're in a position where you're chaining things together or you're wanting things to work in your context or you're wanting to work in a compute constrained environment.
Okay. So vision's pretty unique in that, unlike language and a lot of other places where AI exists, actually vision is almost where humans are not.
Basically like you wanna observe parts of the world where a person isn't present.
Like if we return to our example of like an oil and gas facility where you're monitoring pipelines.
I mean, there's tens of thousands of miles of pipeline You're certainly not gonna have a person stationed every hundred yards along.
It's just an asinine idea and so instead you could have a video theater a visual understanding of maybe key points where you're most likely that have pressure changes.
And to monitor those key points, you know, that you're not necessarily in an internet connected environment.
You're in an operationally intensive environment that even if you did have internet, it might not make sense to stream the video to the cloud.
So basically where you get to is you're probably running something at the edge. because it makes sense to co-locate your compute.
And that's where, like a lot of our customers, for example, use NVIDIA Jetsons.
They're very excited about the digits that was announced at CES to make it so that you can bring these highly capable models to co-locate alongside where their problem kind of exists.
Now, why does that matter? That matters because you can't always have the largest, most general model running in those environments at real time.
I think this is part of, you know, a statement of like the way the world looks today versus how we'll look at 24, 36 and 48 months.
But I do think that over time, even as model capabilities advance and you can get more and more distilled at the edge, There's, I think, always gonna be somewhat of a lag between if I'm operating in an environment where I'm fully compute unbounded, or at least comparatively unbounded in the cloud, versus an environment where I am a bit more compute bounded.
And so that capability gap requires specialization and capability to work best for that domain context problem.
So a lot of RoboFlow users, a lot of customers, and a lot of deployments tend to be in environments like those.
Not all, but certainly some. All right, shift gears here for a moment before we wrap up.
Joseph, you're a multi-time founder, correct?
Yeah. Maybe to kind of set this up, you can just kind of run through a little bit your experience as an entrepreneur.
What was the first company you founded? Well, the very first company was a t-shirt business in high school.
Nice. I don't know that it was founded. There was never an LLC.
Fair enough. My parents knew about it. Uh-oh.
But there is that. In university, I ran a satirical newspaper and sold ads on the ad space for it. and date myself here, but Uber was just rolling out to campuses at that time.
So I had my Uber referral code and I had like free Ubers for a year for like all the number of folks that discovered it.
I kind of joke my first company that maybe the closest thing to a real business beyond these side projects. was a business that I started my last year of university and ran for three years before a larger company acquired it.
And I went to school in Washington, DC. I had interned on Capitol Hill once upon a time.
And I was working at Facebook my last year of university and was brought back to Capitol Hill and realized that like a lot of the technical problem or a lot of the problems, operational problems, that could be solved with technology still existed.
One of those is Congress gets 80 million messages a year.
And interns sort through that mail. And this was, you know, 2015.
So we said, hey, what if we use natural language processing to accelerate the rate at which Congress hears from its constituents? constituents.
And in doing so, we improve the world's most powerful democracy's customer success center.
And so that grew into business that I ran for about three years.
And we had a tight integration with another product that was a CRM for these congressional offices.
And That company called Fireside 21 acquired the business and rolled it out to all of their customers.
That was a bootstrap company. You know, it was nine employees at peak and relatively mission-driven thing that we wanted to build and solve a problem that we knew should be solved, which is improving the efficacy of Congress.
How big is RoboFlow? How many employees?
Well, I tell the team whenever I answer that question, I start with, we've helped a million developers so far.
So that's how big we are team-wise. Team-wise.
Team doesn't necessarily mean, you know, can mean any number of things.
Yeah. Yeah. We're growing quickly. Excellent.
As we're recording this, and this one's going to get out before GTC 2025 coming up in mid-March down in San Jose, as always.
And Joseph, RoboFlow is going to be there?
Yeah, it will be there. I mean, GDC has become the Super Bowl of AI.
Right. Any hints, any teasers you can give of what you'll be showing off?
We have a few announcements of some things that we'll be releasing.
I can give listeners a sneak peek to a couple of them.
One thing that we've been working pretty heavily on is the ability to chain models together, understand their outputs, connect to other systems, And from following our customers, it turns out what we kind of built is a system for building visual agents.
Increasingly, as there's a strong drive around agentic systems, which is more than just a model, it's also memory and action and tool use. and loops.
Users can now create and build and deploy visual agents to monitor a camera feed or process a bunch of images or make sense of any visual input in a very streamlined, straightforward way using our open source tooling in a loginless way.
That's one area that we're excited to show more about soon.
In partnership with NVIDIA and the Inception program, we're actually releasing a couple of new advancements in the research field.
So without giving exactly what those are, I'll give you some parameters of what to expect.
At CBPR in 2023, Robofo released something called RF100, which...
The premise is for computer vision to realize its full potential, the models need to be able to understand novel environments.
So if you think about a scene, you think about maybe people in a restaurant, or you think about a given football game or something like this.
But the world's much bigger than just where people are.
You have documents to understand, you have aerial images, you have things under microscopes, agricultural problems you have galaxies you have digital environments and rf100 which we released is sampling from the Robofill universe, a basket of a hundred data sets, that allows researchers to benchmark, how well does my model do in novel contexts?
And so we released that in 23. And since then, labs like Facebook, Apple, Baidu, Microsoft, Nvidia, Omniverse team have benchmarked on what is possible.
Now, the Rofl universe has grown precipitously since then, as have the types of challenges that people are trying to solve with computer vision.
And so we're ready to show what the next evolution of advancing visual understanding and benchmarking understanding might look like And then a second thing we've been thinking a lot about is the advent of transformers and the ability for... to have really rich pre-trainings allows you to kind of start at the end, so to speak, with a model and its understanding.
But that hasn't fully... made its way as impactfully as it can to vision.
Meaning like, how can you use a lot of the pre-trained capabilities and especially to vision models running on the edge?
And so we've been pretty excited about how do you marry the benefits of pre-trained models, which allow you to generalize better with the benefits of running things real time.
And so actually, this is where NVIDIA and RoboFlow have been able to pair up pretty closely on something that we'll introduce.
And I'll leave it at that for folks to see and tune in to CGC to learn more.
All right. I'm signed up. I'm interested.
Can't wait. So you've done this a few times and, you know, one way or another, I'm sure you'll do it again going forward and, you know, scaled up and all that good stuff.
Lessons learned, advice you can share for founders, for people out there thinking about And whether it's CV related or not, what does it take?
What goes into being a good leader, building a business, taking an idea, seeing it through to a product that you know, serves humans as well as solving a problem.
What wisdom can you drop here? on listeners thinking about their own entrepreneurial pursuits.
One thing that I'll note is you said you'll do it again.
I'm actually very vocal about the fact that Roboflow is the last company that I'll ever need to start.
Like a lifetime's worth of work by itself.
As soon as I said it, I was like, I don't know that.
He doesn't know that. And what if that comes off like RoboFlow is not going to?
I was thinking about, oh, your last company got acquired and so on and so forth.
But that's great though, man. I mean, that's like in and of itself, you know, I suppose that could be turned into something of a motto for aspiring entrepreneurs or what have you.
But that's instructive actually for your question, because I think a lot of people, you know, you should think about the mission and the challenge that you're People say commonly like, oh, you're marrying yourself to you for 10 years.
But I think even that is perhaps too short of a time horizon.
It's what is something that you... A problem space that you can work on excitedly and the world is different as a result of your efforts.
I will also note that, what does it take?
How does it figure it out? I'm still figuring it out myself.
There's like new stuff to learn every single day.
And I can't wait for like every two years when I look back and just sort of cringe at the ways that I did things at that point in time.
But I think that the attributes that allow people to do well in startups, whether they're working in one, starting one, interacting with one, is a deep sense of grit and and diligence and passion. for the thing that you're working on.
Like the world doesn't change by itself and it's also quite malleable place.
And so having the wherewithal and the aptitude and the excitement and vigor to shape the world the way by which one thinks is possible requires a lot of drive and determination.
And so, it's work with people, work in environments, work on problems, where if you have that problem, change with that team and the result that that company that you're working with continues to be realized.
What does that world look like? Does that excite you?
And does it give you the ability to say independently, I would want to day in and day out, give it my best. to ensure and realize the full potential here.
And when you start to think about your time that way of something that is a mission and important and time that you want to enjoy with the team, with the customers, with the problems to be solved. the journey becomes the destination in a lot of ways.
And so that allows you to play infinite games.
It allows you to just be really focused on the key things that matter and delivering customer value and making products people love to use.
And so I think that's fairly universal. Now in terms of specific advice, one things or another, there's a funny paradox of like advice needs to be adjusted to the prior of one situation.
It's almost like the more universally useful the piece of advice is, perhaps like the less novel and insightful it might be.
Right here. I'll note that I pretty regularly learn from those that are a few stages ahead of me and aim to pay that favor forward.
So I'm always happy to be a resource for folks that are building or navigating career decisions or thinking about what to work on and build next.
So I'm pretty findable online and welcome that from listeners.
Fantastic. So let's just go with that segue then.
For folks listening who want to learn more about RoboFlow, want to try RoboFlow, want to hit you up for advice on working at or with a startup, where should they go online?
Company sites, social medias, where can listeners go to learn more?
Roboflow.com is where you can sign up and build a platform.
We have a careers page. If you're generally interested in startups, workatastartup.com is where I see these jobs for. and we've hired a lot of folks from there.
So that's a great resource. I'm accessible online on Twitter of or X at Joseph of Iowa and regularly share a bit about what we're working on.
And I'm very happy to be a resource. If you're in San Francisco and you're listening to this, You might be surprised that sometimes I'll randomly tweet out when we're welcoming folks to come co-work out of our office, like on some Saturdays and some days.
So feel free to reach out. Excellent. Joseph Nelson, RoboFlow.
This is a great conversation. Thank you so much for taking the time.
And, uh, You know, as you well articulated, the work that you and your teams are doing is not only fascinating, but it applies to so much of what we do on the earth, right?
And beyond the earth. So all the best of luck in everything that you and your growing community are doing.
Really appreciate it. Thank you.