Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"Intelligence" itself is very ill-defined and we've never been able to measure it properly, IQ is rife with issues.

At some point, you just have to be pragmatic and measure the questions you want the AI to be good at answering, rather than trying to measure intelligence in general.

In that sense, I see this as one more benchmark that collects questions that we want/expect AI to be good at, is not good at yet and have been underrepresented at previous benchmarks. That's obviously valuable, there's nothing "magical" about it. Although it is reasonable to be annoyed at the "Humanity's Last Exam" naming, of course they must have missed plenty of edge-cases like everyone else and it is very arrogant to claim it will be the "Last" one.



  > "Intelligence" itself is very ill-defined
While this is true, it is well agreed upon (by domain experts) that intelligence is distinct from knowledge recall. But that's what most of these tests... test.

If you look at IQ tests you'll see that they are attempts to test things that aren't knowledge based. You'll also notice that the main critiques of IQ tests are about how they often actually measure knowledge and that there's bias in natural knowledge acquisition. So even the disagreements about the definition of intelligence make clear that knowledge and intelligence are distinct. I feel that often people conflate "intelligence is ill-defined" with "intelligence has no definition." These two are not in opposition. Being ill-defined is more like "I know I left my phone in the house, but I'm not sure where." This is entirely different from "I lost my phone, it is somewhere in California" or "It is somewhere on Earth" and clearly different from "I lost my phone. I'm unsure if I had a phone. What even is a phone?"


Yes agreed, there is indeed a rough consensus on what intelligence is and reasonable ways to approximately measure it. These standard tests have been applied to LLMs from the beginning, they have not proven to be the most helpful to guide research, but there's value to applying benchmarks that have been battle-tested with humans.

It's just that OP was questioning this group's criteria for selecting the questions that determine intelligence. Then we get into endless discussions of semantics.

At the end of the day, you are just testing which questions your AI performs well on, and you can describe how you chose those questions. Claiming it measures "general intelligence" is just unhelpful and frustrating.


They were applied in the beginning because we really weren't that good at solving the tasks. So like any good researchers, we break it down.

But this is like trying to test an elephant but you can't get access to an elephant so you instead train a dog. But putting a dog in an elephant costume doesn't make it an elephant. Sure, dog training will likely mean you can learn to train an elephant faster had you not first trained a dog. Some things transfer, but others don't

I also want to stress that there is a rough consensus. But the ML field (which I'm a part of) often ignores this. I'm not sure why. We should be leveraging the work of others, not trying to start from scratch (unless there's good reason, in which case we must be explicit. But I'm just seeing simple claims of "intelligence is ill-defined" and treating that as if that means no definition instead of fuzzy definition. Which gets extra weird when people talk about moving goal posts. That's how progress works? Especially when exploring into the unknown?)


> IQ is rife with issues

Indeed, and yet people are obsessed with the it and the idea of measuring their own intelligence - I completely do not understand it. I am in an extremely high percentile, but I am a total moron in a lot of areas and if you met me would likely think so as well. It's a poor predictor for just about everything except how good a person is at recognizing patterns (I know there are many different kinds of tests, but inevitably, it feels like this) and how quickly they can reason. But people are obsessed with it (Go on quora and search "IQ", you probably won't half to though, since half the questions there are seemingly about IQ).

A thing I like to say is you didn't earn your intelligence any more than a 7'0" man earned his height - to some degree it seems innate (we don't even really know how).

This all said, it seems even more pointless to try to "IQ" test an AI in this manner. What does it predict? What is it measuring? And you're not going to be able to use the same questions for more than 1 test, because the AI will "learn" the answers.


IQ is a poor predictor of, say, income, in the absolute sense. Correlation is something like 0.4. But compared to what? Compared to personal psychological metrics (of which IQ is one), IQ performs extremely well as a predictor. Things like openness and extraversion correlate at something like 0.1 and others are lower. In fact, IQ is the single best predictor we have and other correlations are usually measured while controlling for IQ.


It’s one of the most studied Quantitative metrics in psychology.


Ok, which iq tests are you talking about? there are like 50 flavors and no consistent school of thought about this, and I hate to be this guy, but can you post your sources here?


What do you consider a poor predictor? It is correlated with many outcomes and performance measures, with increasing predictive power the further you move towards the extremes.

Maybe this is an issue of bubbles, but 90% of the commentary I see about IQ is similar to yours, claiming it is meaningless or low impact.


The lowest IQ thing you can do is be obsessed with IQ.

There are known knowns, there are known unknowns, and there are unknown unknowns. The wise man knows he cannot know what he does not know and that it'd be naive to presume he knows when he cannot know how much he doesn't know. Therefore, only the unintelligent man really _knows_ anything.


IQ is compute speed, not storage. It has nothing to do with knowledge. IBM used to give one out as part of their hiring process, years ago, and even I took it the entire test was a timed multiple choice exam where every question was looking at an object made out of cubes and choosing the correct orientation of the object from the choices, after the object was arbitrarily rotated according to instructions in the question.

Then, IQ can be derived by determining how quickly all participants can answer the questionnaire correctly, and ranking their speeds, and then normalizing the values so 100 is in the middle.

Turns out, scores will fall along a bell curve if you do that. You can call that phenomenon whatever, but most people call it IQ and hopefully I've explained well why that has nothing at all to do with static knowledge in this comment.


  > IQ is compute speed, not storage.
Says who? Honestly. I've never seen that claim before. Sure, tests are timed but that's a proxy for efficiency in extrapolation.

If we define IQ in this way then LLMs far outperform any human. I'm pretty confident this would be true of even more traditional LMs.

Speed really is just the measurement of recall. A doubt we'd call someone intelligent if they memorized the multiplication table up to 100x100. Maybe at first but when we ask them for 126*8358?


> > IQ is compute speed, not storage.

> Says who?

https://en.wikipedia.org/wiki/John_von_Neumann#Mathematical_...

Von Neumann's mathematical fluency, calculation speed, and general problem-solving ability were widely noted by his peers. Paul Halmos called his speed "awe-inspiring." Lothar Wolfgang Nordheim described him as the "fastest mind I ever met". Enrico Fermi told physicist Herbert L. Anderson: "You know, Herb, Johnny can do calculations in his head ten times as fast as I can! And I can do them ten times as fast as you can, Herb, so you can see how impressive Johnny is!" Edward Teller admitted that he "never could keep up with him", and Israel Halperin described trying to keep up as like riding a "tricycle chasing a racing car."

He had an unusual ability to solve novel problems quickly. George Pólya, whose lectures at ETH Zürich von Neumann attended as a student, said, "Johnny was the only student I was ever afraid of. If in the course of a lecture I stated an unsolved problem, the chances were he'd come to me at the end of the lecture with the complete solution scribbled on a slip of paper." When George Dantzig brought von Neumann an unsolved problem in linear programming "as I would to an ordinary mortal", on which there had been no published literature, he was astonished when von Neumann said "Oh, that!", before offhandedly giving a lecture of over an hour, explaining how to solve the problem using the hitherto unconceived theory of duality.

A story about von Neumann's encounter with the famous fly puzzle has entered mathematical folklore. In this puzzle, two bicycles begin 20 miles apart, and each travels toward the other at 10 miles per hour until they collide; meanwhile, a fly travels continuously back and forth between the bicycles at 15 miles per hour until it is squashed in the collision. The questioner asks how far the fly traveled in total; the "trick" for a quick answer is to realize that the fly's individual transits do not matter, only that it has been traveling at 15 miles per hour for one hour. As Eugene Wigner tells it, Max Born posed the riddle to von Neumann. The other scientists to whom he had posed it had laboriously computed the distance, so when von Neumann was immediately ready with the correct answer of 15 miles, Born observed that he must have guessed the trick. "What trick?" von Neumann replied. "All I did was sum the geometric series."


  > Von Neumann's mathematical fluency, calculation speed, and general problem-solving ability were widely noted by his peers
I'm impressed by LeBron's basketball skills. I'm not sure what that has to do with IQ.

Certainly von Neumann's quickness helped him solve problems faster, but I'm not sure what this has to do with the discussion at hand. The story of Polya is not dependent upon von Neumann's speed, but it certainly makes it more impressive. The quote says "unsolved problem." It would be impressive if a solution were handed back in any amount of time.


Isn't that just the Knox cube test, which people with aphantasia are substantially slower to answer? That seems like a very silly hiring test given that aphantasia is not considered a cognitive impairment and people who have it aren't less intelligent in any obvious way.


Speed can be learned though...chess for example.


It isn't really possible to learn calculation speed, learning is about memorising shortcuts and heuristics, or perhaps how to spot them. And training to avoid waste. Strategic questions.

Consider calculating 1+1*1*1*1*1*1*1*1... (= 2). It doesn't matter how quickly someone attempts to multiply an infinite number of 1s they will never succeed because infinity is too large. They have to notice a shortcut that lets them skip doing all the calculations. That shows the difference between calculation speed and happening upon a superior strategy.

But people who can calculate very quickly will have a lot of opportunities to come up with a successful strategy because they can try more in what time they have.


when I was a kid, I had a unique gift for numbers and breaking them down into primes that led me to somehow competing on the county level for these weird speed-math challenges that were popular when I was young. I remember using a lot of tricks. Being able to quickly break down a number into primes by rote memorization was a thing I specifically remember being trained on. There are a lot of number tricks out there that you can train and speed yourself up. this made me quite successful in certain games of chance and odds based games that require quick mental arithmetic when I was younger. Some of it requires mathematical insight, for sure, to derive insight that leads to more speed - but arithmetic can be trained for sure.


This also matches a lot about what we know about the brain and recall. It is a demonstrable phenomena. Just like how any athlete has quicker reflexes. Sure, some might be innate, but their training definitely makes it faster. Information is information.

I mean we all go on "autopilot" at times. Much more frequently in things we do frequently. That's kinda a state of high recall. Not much thinking needs to be done. A great example of this might be a speed cuber, someone who solves rubix cubes fast. Clearly they didn't start that fast.


This is exactly what happens in chess... for instance trading into a known winning endgame.


this was exactly what I thought of and couldn’t articulate it this briefly - particularly pattern recognition and muscle memory in physical chess. It looks crazy when you see it, but the “tricks” are rote. I play mostly 5m chess because I’m a much faster thinker than I am at depth, and a lot of it is just trained speed since I was a young kid. When you see a particular pattern 50,000 times there are people that are good at just immediately making that synapse connect in their head as to the next move, without thinking, I believe this factor is called “intuition” sometimes on accident. It’s definitely a gift to learn though that I think is often confused with deep intelligence sometimes - which explains certain chess types well too. I am often confused for the latter type when I definitely am not - I think those type of intelligences are better at specialization whereas I’m more of a generalist because I can juggle a lot of things at once. They’re both very different types of intelligence that cannot be measured by iq tests, which is why I tend to scoff at what use they are and their usability in predicting outcomes.

Trying to then take this flawed approach and apply it to AI is ludicrous and completely jumping the shark to me. You want to take a flawed measure of human intelligence that we also dont understand fully, and apply it to a machine that we also dont really understand? Ok, then miss me when I laugh at that kind of talk, it is just so silly. This is a more general rant in this broader thread and not directed at anyone spefifically.


I think you would enjoy the book Moonwalking With Einstein. The author is a journalist who's interested in memory competitions and while interviewing he ends up training with these people. Of course, learning that these are skills that surprisingly most people can lean.

I think it's really eye opening into what we can do. The guy trains a year and wins the US competition, moving on to represent the US in a world competition. I think anyone would be impressed if you saw someone memorize a deck of cards in under 2 minutes. But maybe the most astonishing thing is that we are all capable of this but very few can.

https://en.wikipedia.org/wiki/Moonwalking_with_Einstein


> "Intelligence" itself is very ill-defined and we've never been able to measure it properly, IQ is rife with issues.

Yes, because it is 1st person exclusively. If you expand a bit, consider "search efficiency". It's no longer just 1st person, it can be social. And it doesn't hide the search space. Intelligence is partially undefined because it doesn't specify the problem space, it is left blank. But "search efficiency" is more scientific and concrete.


This is always the answer for anyone who thinks LLMs are capable of "intelligence".

It's good at answering questions that its trained on, I would suggest general intelligence are things you didnt want/train the AI to be good at answering.


Are you good at answering questions you are not trained to answer?

How about a middle school test in a language you don’t speak?


For a while I was into a trivia program on my phone. It was kind of easy, so I decided to set the language to Catalan, a language which I never studied. I was still able to do well, because I could figure out the questions more or less from languages I do know and could generalize from them. It would be interesting to know if you could say, train an LLM on examples from Romance languages but specifically exclude Catalan and see if it could do the same.


  > Are you good at answering questions you are not trained to answer?
Yes. Most schooling is designed around this.

Pick a random math textbook. Any will do. Read a chapter. Then move to the homework problems. The typical fashion is that the first few problems are quite similar to the examples in the chapter. Often solvable by substitution and repetition. Middle problems generally require a bit of extrapolation. To connect concepts from previous chapters or courses in ways that likely were not explicitly discussed. This has many forms and frequently includes taking the abstract form to practical (i.e. a word problem). Challenge problems are those that require you to extrapolate the information into new domains. Requiring the connection of many ideas and having to filter information for what is useful and not.

  > How about a middle school test in a language you don’t speak?
A language course often makes this explicitly clear. You are trained to learn the rules of the language. Conjugation is a good example. By learning the structure you can hear new words that you've never heard before and extract information about it even if not exactly. There's a reason you don't just learn vocabulary. It's also assumed that by learning vocabulary you'll naturally learn rules.

Language is a great example in general. We constantly invent new words. It really is not uncommon for someone you know to be be talking to you and in that discussion drop a word they made up on the spot or just make a sound or a gesture. An entirely novel thing yet you will likely understand. Often this is zero-shot (sometimes it might just appear to be zero-shot but actually isn't)


Well ... https://puzzling.stackexchange.com/questions/94326/a-cryptic...

(Someone made a cryptic crossword[1] whose clues and solutions were in the Bahasa Indonesia language, and it was solved by a couple of people who don't speak that language at all.)

[1] These are mostly a UK thing; the crosswords in US newspapers are generally of a different type. In a cryptic crossword, each word is given a clue that typically consists of a definition and some wordplay; there are a whole lot of conventions governing the wordplay. So e.g. the clue "Chooses to smash pots (4)" would lead to the answer OPTS; "chooses" is the definition, "smash pots" is the wordplay, wherein "smash" indicates that what follows should be anagrammed (smashed up).

Disclaimer #1: it took those people a lot more work than it would have taken them to solve an English-language cryptic crossword of similar difficulty, and they needed a bunch of external resources.

(Dis)claimer #2: one of those people was me.

Disclaimer #3: I do not claim that something needs to be able to do this sort of thing in order to be called intelligent. Plenty of intelligent people (including plenty of people more intelligent than me) would also be unable to do it.


Yes — reasonably so, anyway. I don't have to have seen millions of prior examples of exactly the same kind in order to tackle a novel problem in mathematics, say.


Well, LLMs are also remarkably good at generalizing. Look at the datasets, they don't literally train on every conceivable type of question the user might ask, the LLM can adapt just as you can.

The actual challenge towards general intelligence is that LLMs struggle with certain types of questions even if you *do* train it on millions of examples of that type of question. Mostly questions that require complex logical reasoning, although consistent progress is being done in this direction.


  > Well, LLMs are also remarkably good at generalizing. Look at the datasets, they don't literally train on every conceivable type of question the user might ask, the LLM can adapt just as you can.
Proof needed.

I'm serious. We don't have the datasets. But we do know the size of the datasets. And the sizes suggest incredible amounts of information.

Take an estimate of 100 tokens ~= 75 words[0]. What is a trillion tokens? Well, that's 750bn words. There are approximately 450 words on a page[1]. So that's 1.66... bn pages! If we put that in 500 page books, that would come out to 3.33... million books!

Llama 3 has a pretraining size of 15T tokens[2] (this does not include training, so more info added later). So that comes to ~50m books. Then, keep in mind that this data is filtered and deduplicated. Even considering a high failure rate in deduplication, this an unimaginable amount of information.

[0] https://help.openai.com/en/articles/4936856-what-are-tokens-...

[1] https://wordcounter.net/words-per-page

[2] https://ai.meta.com/blog/meta-llama-3/


That’s a very good point. I just speak from my experience of fine-tuning pre-trained models. At least at that stage they can memorize new knowledge, that couldn’t have been in the training data, just by seeing it once during fine-tuning (one epoch), which seems magical. Most instruction-tuning datasets are also remarkably small (very roughly <100K samples). This is only possible if the model has internalized the knowledge quite deeply and generally, such that new knowledge is a tiny gradient update on top of existing expectations.

But yes I see what you mean, they are dumping practically the whole internet at it, it’s not unreasonable to think that it has memorized a massive proportion of common question types the user might come up with, such that minimal generalization is needed.


  > that couldn’t have been in the training data
I'm curious, how do you know this? I'm not doubting, but is it falsifiable?

I also am not going to claim that LLMs only perform recall. They fit functions in a continuous manner. Even if the data is discrete. So they can do more. The question is more about how much more.

Another important point is that out of distribution doesn't mean "not in training". This is sometimes conflated, but if it were true then that's a test set lol. OOD means not belonging to the same distribution. Though that's a bit complicated, especially when dealing with high dimensional data


I agree. It is surprising the degree to which they seem to be able to generalise, though I'd say in my experience the generalisation is very much at the syntax level and doesn't really reflect an underlying 'understanding' of what's being represented by the text — just a very, very good model of what text that represents reality tends to look like.

The commenter below is right that the amount of data involved is ridiculously massive, so I don't think human intuition is well equipped to have a sense of how much these models have seen before.


That's called innovation, something the current AIs aren't capable of.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: