Hacker Newsnew | past | comments | ask | show | jobs | submit | im3w1l's commentslogin

A very simple program that loops over all strings and feeds them into a proof verifier should eventually prove every statement that can be proven, as far as I can tell?

No proof verifier verifies all valid proofs and terminates on all invalid proofs.

(actually I am wrong. You would introduce a new proof, and then step the verifier on all ongoing proofs, so non-termination isn't a driving concern)


No. Whenever you are done with all strings of length n, you still have to check all strings of length n+1. So that moment you identify by "eventually" can never be reached.

If something can be proven there is a finite length proof, and if you check lengths one by one eventually you will reach a high enough length for a valid proof.

No, but you hit the nail on the head, that’s the most interesting part.

The proof verifier uses fixed math axioms. The busy beaver function at high enough N cannot be proven with those axioms.


I specifically said that it will prove everything that can be proven, conceding that some statements cannot be proven no matter what you do.

The point is that some statements could be provable, but not with today's proof verifier.

"Everything that can be proven" is relative. PA can prove some things, ZF more things. In 200 years we could develop more powerful math foundations which can prove more things. Today's proof verifiers could never prove them, but tomorrow's proof verifiers could. And the cycle repeats.


Just like turing machines are universal in the sense that they can all emulate each other, I feel like there should be some universal logic that can emulate any other logical theory. Something like the gödel numbering construction maybe? This is where my knowledge ends, I'm afraid.

Nope, there is no such universal logic. Godel helped show the opposite actually (incompleteness). I think Scott Aaronson's post is very fascinating explanation of this stuff. https://www.scottaaronson.com/papers/bb.pdf

I don't think incompleteness disproves my idea, at least not trivially. Let's take the case of ZF vs ZFC. I would say that ZF can simulate ZFC. To create a simulation we want a function f from statements in ZFC to statements in some subset of ZF so that valid inferences in one correspond to valid inferences in the other.

This is quite simple. f(p) = C implies p does the job quite elegantly.

Interestingly it's harder to do the opposite, to simulate ZF in ZFC, because there is no way to express "forget that you know about C". Such a construction cannot be possible in general because if a contradictory axiom is added, then everything is true, and a theory where everything is true is useless and can't simulate anything. However for C in particular I believe it should be possible to make such a construction but I can't immediately think of how I would do it.


Personally I think the dirty secret of the brain is that a lot of things are hard coded. And many things that we need to learn are also hard coded except that some parameters need to be tuned.

If we puke, the brain will not do general aversive learning, it will learn to avoid specifically the last thing eaten, because it instinctively knows about food poisoning.

Imprinting is absolutely fascinating. Some newborn animals will run a very simple pattern detector like looking for a red dot or something and use that to bootstrap their conception of their parent.

For fully general learning I have a hunch that it can be done using local history plus a semi-global reward scalar (global neurotransmittor levels).


regardless if the intelligence in the brain is hardcoded or not, to the extent it is, this information must have been compressed in the genome, which runs counter to almost all observations: a child doesn't remember the experience of their ancestors, for example. The only sense in which we do carry mental state without relearning is emotions, instincts, reflexes (some neuronal pathways that connect the eye to the middle ear), hormonal driven behavior (fear adrenalin).

For another, there are about 200k promotor regions (including non-coding) in the human genome.

A promotor region might have say 6 to 15 bits of information.

Can you compress 2025 or even 2024 era LLM intelligence into 3 megabit = ~400 kB ? I think not. I think a lot of compression is still possible, but 400 kB?

So I think we can box up the idea of "dirty secrets of the braing: not learning but hard coding". There is a lot of hard coding in biology, but brains are evolved specifically to enable learning within the individual lifetime instead of only learning by natural selection.

I also don't buy the following argument:

> If we puke, the brain will not do general aversive learning, it will learn to avoid specifically the last thing eaten, because it instinctively knows about food poisoning.

Each time it happens that I end up puking, I do feel aversion and try to avoid puking at all, sometimes I succeed but sometimes is just puke. There must be fundamental puke reflexes (which one fails to avoid) and avertable puke reflexes.


You are thinking on the wrong level. Of course we don't have an encyclopedic knowledge of the world encoded into our genome. It's like we have certain structures of the world hard coded, and they may use different learning algorithms.

We instinctively know that there are other intelligent beings, and we have the the machinery to model them. We are born with the capacity for language. It must still be learned, the specific words aren't hardcoded, but the concept of language is. We are born with the capacity to store and replay memories. The brain knows some aspects of how the world is supposed to look like visually, and if it doesn't it will try to correct that. People who used optics to see the world upside down have found that after a brief time their brain learned to flip the world right side up.


> Can you compress 2025 or even 2024 era LLM intelligence into 3 megabit = ~400 kB ? I think not. I think a lot of compression is still possible, but 400 kB?

There are a few extra levels of interpretation (like protein synthesis) that are more like a transpiler than compression (imo), over a 4-base language that is read in a sliding window and is affected by surrounding conditions, so the same "token" sequence may produce different things depending on external factors. Some biologists I used to collaborate with talked about 7 layers to this process, I have only described one level here


From an information theory perspective it does not matter how many levels of interpretation are in between, that hardcoded information must pass the genome. for example we can discuss "what about protein synthesis", for example binding affinities, folding helpers etc. they in turn were encoded genetically as well.

> There are a few extra levels of interpretation (like protein synthesis) that are more like a transpiler than compression (imo), over a 4-base language that is read in a sliding window and is affected by surrounding conditions, so the same "token" sequence may produce different things depending on external factors. Some biologists I used to collaborate with talked about 7 layers to this process, I have only described one level here

so the same "token" sequence may produce different things depending on external factors.

yes, non-hereditary learning depends on external factors, thank you for paraphrasing me while shifting attention.

the multi-scale nature (transpilers etc.) doesn't change the theorems in probability and information theory which seriously constrain the maximum amount of information a message can store.

The whole point of a brain is that it is an organ dedicated to storing, retrieving and timely utilising information one can't afford to store in a genome.


I'm not convinced information theory is the right avenue for such non-deterministic, environmentally affected, open ended systems. Protein synthesis is defined by far more than the DNA, which is always being processed, most of which is junk and being recycled. The layers in between are more than interpretation of the prior layers, the data grows at each level to incorporate more sources under your information theoretic formulation.

I was never paraphrasing you, but thank you for attempting rhetorical antics?


Information theory is the right avenue but it's too often used to imply things it says nothing about. From (1)

1. A determinist transformation can't create more information than its inputs

you can't derive (2)

2. Later stages don't contain any new information.

without restraining your model to a monotonic chain of transformations that can't tap into the information content of the environment.


> I'm not convinced information theory...

I'm not convinced my molecules follow the laws of thermodynamics.


It kinda rubbed me the wrong way that he considered both the lost sale and the value of the goods as part of the cost. That's double counting.

its an ai written article it seems. probably written for engagement bait.

It does indeed mean ideological alignment. But we don't get to leave the answer blank. They have to pick an ideology to put in there, and whatever they pick will have huge consequences.

My point is maybe we don't let some of the world's richest people with some very... interesting ideas pick which ideology we digitalize? Maybe we figure out a way to give people a say in this?

It's kinda a throw-away remark in the middle, but I thought it was really cool to realize that speakers produce an on-axis net-blow and an off-axis net-suck.

For public transit in my city they have an official website where you can type in where you are and where you want to go and it plan it all for you. Can choose from arrive asap to picking a time when you want to depart or a time you want to arrive. This has been operational since 2002.

More recently Google Maps has the same functionality and it will work in many cities across the world.


I think their broader point is that it's becoming common to reach for an AI chat to solve problems that should take a few seconds of thinking to solve. Notably this kind of thing is not a search problem, since you presumably already know the schedules for both busses and just need to come up with a reasonable allowance for schedules not lining up properly.


Does that mean you can build precedents with "matchfixing"? Like pay the plaintiff under the table to throw his case by presenting really bad arguments? And then subsequent cases must reference that result?


You are making an argument against caching it at all which is clearly not what they want. So the comparison must be uncompressed storage vs compressed storage.

The compression cost is always the same. The storage cost depends on how long you keep something in cache. The decode cost is proportional to number of total hits.

So compression makes the most sense for something you want to keep a long time that will be accessed very rarely. And the least sense for something you will drop very soon and will have many people requesting it.


Maybe cloudflares workload is substantially different, but i think youre missing just how long that tail was. A “lot”, maybe half, of unique objects werent requested a second time in any meaningful period time. Like days. And edge nodes have nowhere near the iops or cycles to spend doing _any_ extra work. So yes it is a waste of resources to cache or process in any way.


Well what do you want? Presenting clear, desirable, and achievable visions and trying to build consensus for how AI should develop is crucial at this point in time.


> Ban construction (...) and let the economy and market forces compete to find the cheapest alternative

This sounds nice and unbiased on paper, but often this means that rather than expanding alternatives, the poor are simply made to go without and then they will scream at politicians for relief until they get it and you are right back at square one.

As a fellow Swede you will know how the price of gasoline has become a hot topic and how politicians campaigned on bringing the price down again by reversing climate policies.

Personally I think the solution is to make a commie-style 5 year plan and then work backwards to what kind of incentives would be needed for that plan, probably involving industry leaders in that discussion. Then you can make sure that the incentives are technology neutral. This ensures that if nothing better comes along the "free" market will execute 5 year plan but they do have the option of gambling on alternatives if they so choose or if some wonder technology emerges out of a secret lab.


A commie-style 5 year plan or even 10 year plan would be kind of great. Electricity is one of those products that we know what we want long term and should be willing to budget for it in a reasonable way. We need to double the production while eliminating emissions, while at the same time restore rivers and prepare for more extreme weather. We also need a stable energy market without shortages, as week/month long shortages makes citizens and businesses go to the government for solutions. As the last energy crisis illustrated, people can't just stop going to work because the market price for electricity is too high. People need to eat and pay rent, and those needs is why they go and scream at politicians to fix it.

The price of gasoline is a actually a good example of an fairly open debate over what the costs should be on a national level for a global product. We are currently paying one of the highest electricity prices in EU, not because the energy is expensive to produce at the wind farms, but because construction of new transmission from the place of production to the place of consumption is carried by the consumer. There is very little political discussion on what a fair price is for electricity, and how the cost of the grid should be fairly paid.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: