Today on the podcast, we're talking about Stack Overflow, which is essentially recreating itself into an AI data provider.
I think the reason I want to cover this is because I think we're going to see this exact same trend played out with a ton of different online companies that are struggling with lower web views, lower usage, after ChatGPT and a lot of these other AI tools came out that will answer questions faster for you.
Stack Overflow is one that has been reported on extensively and seen a dramatic drop in usage.
But you can also talk about Wikipedia, you can talk about Chegg, you can talk about a lot of different companies that would you know, do kind of questions and answers in specific niche areas.
The AI models came in, scraped their whole website, have all of that baked into their.
The original companies are suffering because no one really is using them.
So we're going to get into the future of some of these forum type websites and specifically the deal that Stack Overflow has done, how it's similar to other players and more on the podcast.
Before we do, I just wanted to mention if you want to try any of the AI models that I talk about on the show.
I'd love for you to try out my startup, which is called AIboxai.
You get access to the top 40 different AI models Google Gemini OpenAI Anthropic Cohere DeepSeek, everything image models, audio models, like 11 Labs all for 20 bucks a month, in one place, on one platform.
So if you don't want to have to get a new subscription every time you want to test out a new AI tool, go check out AIboxai.
There is a link in the description.
All right, let's talk about Stack Overflow.
I think this all came out at Microsoft's Ignite conference.
So Stack Overflow came out and showed a whole bunch of new products that they were going to try to essentially use to position themselves as a really useful part of the enterprise.
AI stack right.
Every enterprise needs to have a license to this new Stack Overflow tool.
This is kind of a new take for the company.
Stack Overflow definitely struggled after ChatGPT came out.
And there's a number of articles that just said their web traffic went down significantly, right?
This is traditionally a website where developers would go on and ask coding questions and say, hey look, I'm running into this issue.
Does anyone know how to fix this bug in my code?
People would respond and help debug or work on code problems together.
And you saw this play out in a lot of different industries.
I mentioned Chegg, which was like for students.
Students would ask questions and other students would respond.
So it's kind of like more like an education side.
Of course, we saw this with Wikipedia, who has recently said that they are seeing a massive drop.
I don't want to say massive, but they are seeing a decline in web traffic that is from humans and an increase in web traffic that is from AI scrapers bots, and maybe even some of those are agents.
And Wikipedia has responded by making an API.
Chegg has been struggling.
And then we also see companies like Reddit, who again is a forum, but like for everything.
And Reddit has went ahead and made these deals where they'll license their content to companies and they're able to just make kind of like blanket deals.
So with all of that, a lot of the new tools that they're making are specifically at Stack Overflow.
They're specifically designed to feed into internal AI agents that are using the MCP or the model context protocol.
And they're using that with different variations designed specifically for Stack Overflow.
It's essentially, you know, Stack Overflow internal is what the new tool is called.
And it is essentially an enterprise version of the web forum that they have.
But they have a bunch of additional like security and admin controls on it.
So companies have that extra security and control over the content.
This is their CEO.
Parashna said talking about all of this, said that they were already seeing a whole bunch of enterprise companies using their API for training.
So that's another thing that Stack Overflow did, right?
They saw kind of like Wikipedia, they're like, look, our traffic has dropped a lot.
And we have a lot of, you know, bots that have been scraping us, they just made an API.
And they're like if you're an AI company, you should use our API for training or you have to.
It's our term of service, otherwise we're.
They saw a lot of.
Apparently, according to their CEO, they saw a lot of success and progress with that specific model.
And so then they decided to kind of take this new product direction where they're like.
Well, maybe enterprises would want access to Stack Overflow tied directly into the AI models that they use in a very direct way.
They already made a bunch of different content deals with a whole bunch of AI labs that essentially allow them to train their models on public Stack Overflow data.
And they're doing this just for a blanket fee.
So it's very similar to the reddit deal which happened, and the reddit deal has brought in more than 200 million dollars for reddit.
Just, you know kind of giving like i think reddit is working with open ai and google specifically that i know of, and i think it's like 100 million dollars a piece.
They're like look, you can scrape reddit 100 million bucks and um you, you can kind of have access to this.
So it's like a big boost in revenue for the company.
And, of course, open ai and google are like well, we don't have to deal with any lawsuits.
It's a great data set.
Um and others are kind of blocked from it, so it made sense for them.
A really important part of this new product is a layer of metadata that Stack Overflow has access to.
Because you could say well, if all of their website has already been scraped by the AI models, why does everyone want access to maybe having this custom API into it?
And the reason why is because they still have some data that others don't have.
Beside the questions and answers that you see inside of Stack Overflow, the data also includes some information like who answered the question and when they answered the question.
They also have content tags and a lot more complex assessments of some of the internal coherence.
So what this means is you could say like, look, I'm asking a question about Java.
But you know, this question was answered like back in 2012.
So is it relevant to the current version of the coding language I'm running today?
Or maybe I'm running, you know, I'm using an old version of some coding language or some tool and I need, like an older answer.
And so what's interesting here is, because they have that date, not a lot of these AI models scraped that.
And so they're actually able to assign this sort of like an assessment score to say how likely the it's a reliability score which will tell the AI agent how likely the answer is to be trusted right.
It's like well, based off of what you're currently asking the question about your current stack and when this answer was created, this is how likely it is to be good.
And in addition to this, they know who answered the question.
So they're actually able to look at those accounts and see, you know how legitimate the accounts are, how you know how good of a developer they are, how good their solutions are.
And then they can use all of the data from the individual users or contributors account to determine how good the answer will be.
So this is interesting, right?
And I really appreciate this.
They're trying to lean on some data that other people might not have, that they have exclusive access to, and make the product better.
The CTO, Jody Bailey, said this about it.
They said the customer can set up their own tagging system or we can dynamically create that for them.
What we'll be doing in the future is really leveraging that knowledge graph to connect people and to connect concepts and pieces of information, rather than requiring the AI system to do that on their own.
So while Stack Overflow right now is making a whole bunch of tools for enterprise agents, it isn't building all of those agents itself.
So it's kind of hard to say what their final product is actually going to look like when it rolls out.
Bailey is really excited about the writing function, though.
Bailey is their CTO.
And Bailey said that the writing function is going to allow agents to create their own Stack Overflow questions.
If they can't answer a specific question or they notice there's like a knowledge gap, they're actually able to ask a question on Stack Overflow.
I think my question is will real humans seeing AI bots ask questions on Stack Overflow feel obligated to answer a bot right?
It's not really like a human, but there's usually a human behind the bot asking the question.
So maybe they'll still be helpful.
I'm not sure.
Or are they going to just have AI bots come in and try new ways to answer the question?
It's going to be interesting.
The way that Bailey sees it right now.
This kind of like read write function means that, as the quote is, as we continue to evolve, it will require less and less effort from developers to capture the unique information about the way they operate their business.
So Overall, I think this is a fantastic direction for Stack Overflow.
They're leveraging pieces of the data that only they have access to.
And I think we're going to see a lot of other companies that have these kind of question and answer forums, which are essentially deep sources of data.
We'll have to monetize it one way or another.
The blanket deals are one thing, but I think it's great if they're actually building tools and software that people can use and add extra context and data that the scrapers don't have access to.
All right.
Thank you so much for tuning into the podcast today.
If you enjoyed the episode, make sure to leave us a rating and review wherever you get your episodes, and make sure to check out AI box dot, AI for all of the best AI models in one place, on one platform a month.
There's a link in the description.
I will catch you guys all in the next episode.