Fascinating project.
How are the probabilities being measured? I’m skeptical of AI analysis because my past experience is that it appears to be analyzing, but behind the curtain there’s not even a fraction of the same rigor a human analyst would have done. In other words It’s very shallow but looks deep.
In this model the probabilities were measured by an ensemble of LLM's. We give them context and ask them to follow a specific methodology for making a forecast, and then aggregate their forecasts into a "crowd" forecast. We can get input from humans as well to compare/combine them further, but in this particular instance I just used our AI Forecaster.
You're right that a human analyst - especially a very experienced one - is going to be able to still do a better job of analysis most of the time, but the AI is useful for having loops to re-examine the question, consume and curate new information, etc. We've built the system with that assumption: let humans do what they're best at, let AI do what it's best at...
I think this comment encapsulates the fundamental state of AI at the moment.
The industry believes very strongly that this simulation of methodology is basically as good as the real thing (or at least "good enough" in most cases).
This belief is the fault line between the bullish and the bearish; it requires a leap of faith that not everyone is willing or capable of taking.
Every forecast question we ask has a verifiable outcome. We then score it based on what actually happens in reality and create a track record for all of it. In typical analysis, nothing is scored, and no one goes back and checks if it was right or wrong, but we do - whether we're scoring humans in human forecasting exercises, or seeing how AI does.
The kicker has a drug habit. The quarterback is making a suspicious number of overseas call. The coach is having an extramarital affair that could be exploited by a blackmailer.
A rival team has an ambitious but craven assistant coach who, with a regime change, could be better controlled towards our interests.
Another team's key player couldn't be replaced on short notice, were they to be removed from the board, in a deniable 'gang-related' incident in front of the nightclub they frequent.
The roster of a favored team could be boosted with secret financial incentives to candidates, funded off-the-books by new gambling and prostitution operations.
Small potatoes- more realistic: Les Wexner pumped tens of millions into the university to whitewash his reputation; keep them perennial favorites but grease some SEC refs as necessary to keep our guy out of the limelight
Les Wexner pumped tens of millions into the university to whitewash his reputation;
The Henkel-Sika-3M lab for sustainable adhesives (I'm just making that up, one doesn't exist) would have a different name and a whole lot less funding if it didn't produce research indicating that whatever it's benefactors want to do is in fact sustainable.
I grew up with a kid whose dad worked in security for the St. Louis Rams NFL team. I was shocked when he told me how many attempts there are to blackmail or otherwise extort money from players.
This could be a whole sub plot in a comedy with a bunch of out of touch analysts. "he's with her because he's into thick ones", "no that's just how black guys are", "no she must be the throat goat". You could even add a whole additional layer of jokes into it where they project snobbish DC norms onto people who aren't that.
Unfortunately nobody will ever make a serious workplace comedy that ridicules federal intelligence because hollywood and the feds are buddy buddy.
I find it really cringey when people describe a whole intelligent system… and it’s just asking copilot to generate slop.
It didn’t “calculate probabilities”, it just hallucinated slop. The same kind of slop almost started WW3 when a similar “intelligence analyst” system told the US navy to attack China because they were smuggling nukes into Iran. They weren’t, it was all slop. But their system was full of “calculated probabilities” too!
Not exactly - all the forecast questions get scored so we know how accurate the AI is and how calibrated it is. It's no different than asking a large crowd of humans on Good Judgment Open or Metaculus these same types of questions. Everything we ask eventually happens or it doesn't, and we can score it all.
I really enjoyed reading this, got invested in learning about this idea, and then realized they they never actually show any evidence that this works? Like the text-based model they are experimenting with did a kinda counterintuitive thing that they thought was worth a section, and then they don't say whether or not it was right? The idea of aggregating chatbot context into a probabilities is super interesting and I really want to know if it's bullshit or not
I think that was the entire premise of the CIA analyst framing - you don't know which one of several outcomes is going to prevail but you have a model which constantly updates its priors to increase the probability of one outcome over others.
In this case, you still don't know whether some guy's passing percentage stabilizing over a baseline will help you win the big prize or not but it does tell you that he may win some sort of recognition for being a standout player or that the guy who scouted him deserves credit.
Fascinating project. How are the probabilities being measured? I’m skeptical of AI analysis because my past experience is that it appears to be analyzing, but behind the curtain there’s not even a fraction of the same rigor a human analyst would have done. In other words It’s very shallow but looks deep.
In this model the probabilities were measured by an ensemble of LLM's. We give them context and ask them to follow a specific methodology for making a forecast, and then aggregate their forecasts into a "crowd" forecast. We can get input from humans as well to compare/combine them further, but in this particular instance I just used our AI Forecaster.
You're right that a human analyst - especially a very experienced one - is going to be able to still do a better job of analysis most of the time, but the AI is useful for having loops to re-examine the question, consume and curate new information, etc. We've built the system with that assumption: let humans do what they're best at, let AI do what it's best at...
It seems very likely then that the results are hallucinations. You can’t prompt an LLM to follow a certain analytic methodology, it can’t.
But, it can simulate the output of a methodology, which is not the same thing.
I think this comment encapsulates the fundamental state of AI at the moment.
The industry believes very strongly that this simulation of methodology is basically as good as the real thing (or at least "good enough" in most cases).
This belief is the fault line between the bullish and the bearish; it requires a leap of faith that not everyone is willing or capable of taking.
Every forecast question we ask has a verifiable outcome. We then score it based on what actually happens in reality and create a track record for all of it. In typical analysis, nothing is scored, and no one goes back and checks if it was right or wrong, but we do - whether we're scoring humans in human forecasting exercises, or seeing how AI does.
The kicker has a drug habit. The quarterback is making a suspicious number of overseas call. The coach is having an extramarital affair that could be exploited by a blackmailer.
A rival team has an ambitious but craven assistant coach who, with a regime change, could be better controlled towards our interests.
Another team's key player couldn't be replaced on short notice, were they to be removed from the board, in a deniable 'gang-related' incident in front of the nightclub they frequent.
The roster of a favored team could be boosted with secret financial incentives to candidates, funded off-the-books by new gambling and prostitution operations.
Nah.
These people don't like to share. They have to pay and/or payoff players, coaches, staff and so on.
But after that?
Yeah.. they're taking every nickel they can right down to the bottom line.
They're not letting anybody else in on that racket.
Small potatoes- more realistic: Les Wexner pumped tens of millions into the university to whitewash his reputation; keep them perennial favorites but grease some SEC refs as necessary to keep our guy out of the limelight
The Henkel-Sika-3M lab for sustainable adhesives (I'm just making that up, one doesn't exist) would have a different name and a whole lot less funding if it didn't produce research indicating that whatever it's benefactors want to do is in fact sustainable.
The radio station chief is a con artist rainmaking the agency, he will cut us in if we squeeze him hard enough.
I grew up with a kid whose dad worked in security for the St. Louis Rams NFL team. I was shocked when he told me how many attempts there are to blackmail or otherwise extort money from players.
Ah dang, this looks like it got hugged to death.
He can't hit a curve ball... And an ugly girlfriend.
Ugly girlfriend means no confidence.
This could be a whole sub plot in a comedy with a bunch of out of touch analysts. "he's with her because he's into thick ones", "no that's just how black guys are", "no she must be the throat goat". You could even add a whole additional layer of jokes into it where they project snobbish DC norms onto people who aren't that.
Unfortunately nobody will ever make a serious workplace comedy that ridicules federal intelligence because hollywood and the feds are buddy buddy.
Suspect you suck even more joy out of it the further removed from original local friends and neighbors getting together it started out as.
The Culinary Institute of America doesn't have a football team that I can find, but I would still like to analyze it.
I bet the CIA could field a half decent MAC team.
I find it really cringey when people describe a whole intelligent system… and it’s just asking copilot to generate slop.
It didn’t “calculate probabilities”, it just hallucinated slop. The same kind of slop almost started WW3 when a similar “intelligence analyst” system told the US navy to attack China because they were smuggling nukes into Iran. They weren’t, it was all slop. But their system was full of “calculated probabilities” too!
Not exactly - all the forecast questions get scored so we know how accurate the AI is and how calibrated it is. It's no different than asking a large crowd of humans on Good Judgment Open or Metaculus these same types of questions. Everything we ask eventually happens or it doesn't, and we can score it all.
Okay. So how accurate is it?
Ok, so in an ML system you can calculate things like cross entropy loss, which penalizes your model for making confident, inaccurate predictions.
It doesn’t care whether your system is made of LLMs, decision trees, or bananas.
However, the problem is that unlike something like a neural net, you can’t exactly use backprop to improve.
I really enjoyed reading this, got invested in learning about this idea, and then realized they they never actually show any evidence that this works? Like the text-based model they are experimenting with did a kinda counterintuitive thing that they thought was worth a section, and then they don't say whether or not it was right? The idea of aggregating chatbot context into a probabilities is super interesting and I really want to know if it's bullshit or not
I think that was the entire premise of the CIA analyst framing - you don't know which one of several outcomes is going to prevail but you have a model which constantly updates its priors to increase the probability of one outcome over others.
In this case, you still don't know whether some guy's passing percentage stabilizing over a baseline will help you win the big prize or not but it does tell you that he may win some sort of recognition for being a standout player or that the guy who scouted him deserves credit.