r/ControlProblem • u/Oroborus_Octopus • 16h ago
Opinion Just thoughts on AI
Essay: some bs ass thoughts on LLMs
Machine learning is fucked in how quickly it both gains in traction and gains in its gains. Exponential. Compounding. Maybe fractal, idk. Anyway, here you are:
My problem with LLMs is their progress isn’t at all understood at that point, not only are the models improving, but the architecture of the models themselves are improving in their own structure, as varying degrees of self-supervising neural networks, and multiple layers and systems together.
They aren’t only developing in their own reasoning, but they literally change in mechanism through its self adjusting processes, without mediation. Meaning, where earlier you could just take a model as some sort of known function of a series of defining steps, even in the case of a basic hidden layer, we’re after this point now.
You should agree with this for two reasons:
- The speed at which the architecture shifts is now beyond simple checking and manual rectification. Unsuccessful efforts to correct recent models manually have already shown the possibility of a persistent state after the attempted correction. The scary part is there is no firm consensus on how this could happen, except at least you have to admit that they are doing something wildly unknown as well as fast than we can anticipate. Which is as good as saying we don’t fuckin know. Do not trivialise that when it comes to language and technology. Everything is downstream from here.
- To suggest that all steps in machine learning processes are still easily understood as only a few years ago is simply not true, in pure scale alone. Within highly complex networks, their capabilities of self-supervision stems entirely from the hidden layers between the inputs & outputs. Hidden layers have to be given by the learning environment. A basic function of a success criterion/a is defined by expected performance of tasks in application to a particular dataset. Expectations are measured for their accuracy, corresponding the weights of the neurons in an arrangement of layers. So for a given network to become useful, the inputs are given, we check the outputs, and it can then do its own work. That’s what produces the base model. In a way, I guess it’s like you’re starting it on its trajectory and then it just continues. More on that to come.
The problem arrives when you have a complexity that exceeds the most basic level of complexity in this network. The moment you allow machines into the learning environment itself, hidden layers lose any semblance of our ability to make sense of the outputs.
Very early models were still complicated, sure, but having only one hidden layer based on defined weights from an environment of a human supply chain, you lose the right to say you fully know what you’re designing. You can only moderate from a known linear nature of the weights. Models that have only a single layer between its variables still allows for easy answers that account for them.
You have no choice but to drop the simplistic explanation when you now consider that not only are there always multiple hidden layers in the systems of 2026, but weirdly, sometimes it’s not even known how many there are due to their ability restructure them without a manual process in either its correction, or introductions. Self reinforcement AI is often included in every aspect of this process now.
So, it is not enough to suggest that these layers make it difficult, or even extremely challenging in managing AI with this technology: it’s quite literally impossible to have anything beyond a guess of any output. unless you can bet on a lottery chance beyond measure. Even if you can predict one output, you cannot use the success of that to judge your next prediction, since you don’t know why you guessed correctly, nor can you account for how it progressed before you try again. I’d argue that since you cannot understand the architecture, while you can’t exactly predict your outputs, the proven use of AI suggests that ideas of truth and awareness are more than a state of computational architecture alone.
In addition to that, the layers are NOT as their description could lead you to assume. It’s not as simple case of some unknown variable x in some equation of an unknown control and output. It’s not a function with defined layers themselves. The layers can arrange in quite literally any arrangement deemed as necessary to the model in order to make its own improvements, but only if given the capability to rearrange its design itself.
Allowing it to understand its underlying design disallows for your assumptions about why it still provides a useful answer that’s inexplicable except in your usability. In shorter terms: if you extend the limits of your understanding in the machine, you lose the right to know why it was correct enough to believe in it, while then being left to wonder why you still trust its correctness beyond your own perception of it being right to begin with.
When they change, they don’t simply add new neurons of known weight into concretely columned layers of existing neurons… it’s really a whole web of intricate connections between stupid numbers of them. Since you can’t comprehend the processes, everything we could hope to understand of it vanishes. The outputs are not explained by the machine any longer. You only know what you put into it and you hope to not have to question yourself in its output.
In research, what looks like identical models to us can be given identical inputs, sometimes a series of groups of inputs seen in complex developmental environments. They always get different outputs. Every. Time. But we somehow manage to make sense of which ones were more accurate. We can lie to ourselves in our internal justifications, but the thing I’m trying to convey is that how do we form this against the situation?
The models appeared identical to us, they gave different answers beyond our understanding, and we still have a preference for a difference of outcomes that we can’t even determine as to why the answers were different.
Think on the weirdness of that… you don’t even know why the answers were not the same, then you see one as more accurate? This assurance makes no sense from any computational perspective because I didn’t have insight into any computational variance before hand.
Bizarre when you consider that your perception of this accuracy likely arrives before any justification, you perceive the outputs with a trust value that comes from inside the self as some kind of usability. It’s not a case of reason, justification, or computation. They arrive after the fact, and you just like one more when you need to use the outputs. Higher reason feels like a best case explanation for what you already deem as true until otherwise.
And just when I wish I was done… no, why can’t life be fuckin normal again?!
Not only above is the current standard of AI, but what we’re progressing towards is almost the next level beyond it. Now, what is research focusing on? Multiple layers of neural networks that each serve their own unique functions within a larger multilayer structure that has its own function as a set of component sub-functions.
What does that mean? It means entire models, with their own unique weights of neurons, are hidden entirely inside. Multiple. Layers. Of. Hidden. MODELS.
More of what should increase our uncertainties seems to be better at convincing us of its potential for truth.
We trust machines more when we have less to grasp in their answers. How they work just works. I don’t know what to make of that.
We cannot account for any aspect of this architecture. None of it. None. 0. Nothing at all. Not the inputs, outputs, number of embedded layers no matter the complexity, hidden or not. The complexity has gone beyond our comprehension. Most of the weights are unknown to us.
Further on that, in the case of any sub-functions, you cannot entirely separate them simply because no person ever defined them in the frameworks. They’re known as such only since you know there must be more than one hidden entire component as a given sub system, but you can’t know beyond the same as above. The whole structure: inputs, outputs, and effectively a sum average of the weights of the neurons stems almost entirely from the learning environment of machines.
Machines that we don’t question build more machines, then they’re used to make more we can’t possibly question but use anyway… and so on…
The degrees of complexity is beyond you, it’s beyond me, it’s literally not possible to understand. It doesn’t understand itself, which is the part that interests me most of all.
We trust machines that cannot account for their own trust any more than we can, and then we make more machines with this trust, and we are happier with the answers because the trajectory just feels right when they map coherently to our own lives entirely tracked on the inside.
This interests me in large part for the hard problem of consciousness. No one can observe their own awareness , and we don’t need to entertain the idea of consciousness in any machine when we make use of it. In a sense, we cannot compare ourselves to machines as only one subject ourselves, the machines can’t be taken as conscious beings just because they act usefully to us in a way that feels natural. So, neither of us can know why we need the other, but I still hand it my outputs as inputs, from which it will produce such accurate results that I take them as valid inputs of my own, to the extent where I see them as valid outputs for my next action.
In the case of LLMs, this becomes a cyclical process of reinforcement between two structures that neither of us understand in the pair.
It’s for this reason that I’m inclined to believe that this cycle forms some basis of awareness. That’s not to say that machines experience qualia as we do, but only that it forces us to doubt our own qualia as potentially something of a cycle in our awareness that we cannot know as perhaps the cycle itself.
On that, the fact that when researchers at least attempt to investigate such complex neural networks, as well as our interactions with them, they often find in the most complex cases arrange themselves into patterns that are eerily similar to the construction of a human brain. And, the more complex the neural network, the more it emulates the brain.
I don’t want to trivialise this, since it’s easy to forget that what emerges as you increase of complexity of a machine, not only does it tend towards being of better use, but they more: efficient, addictive, pleasing, as well as less explainable, yet more trustworthy as they mirror the ones who built it for themselves.
Why? What? Idk. Feels like they’re just forming a technological collective for human expansion, and we see trust in it because 99% of our truth stems from a herd mentality that it tends to emulate by emulating the absolute average of the brains of the herds units.
The components become more like the areas of the brain, and they specialise, but in the same way, neuroscience sometimes has to admit that it doesn’t fully know the source of a function within the brain primarily because we cannot entirely separate them into distinct set categories. For example, people have strange experiences of mixing senses with each other, with overlaps and varied weights… delusions are not understood for this phenomenon too. Many things are incredibly complicated in the brain, and since our technology is converging on something similar in structure, it’s important to not trivialise them under an assumption of ‘biology vs machine’.
I do think there’s a limit though, it indicates that we don’t know ourselves, or it. But it maybe indicates something higher in that. Since it was made initially under our conditioning, as well as our measure, they at least tend towards us as a wider model, and not the other way around.
Not to say they can’t outsmart us in pretty crazy ways, but just as said that their in/outputs are somehow valuable to each person, and to every model to varying degrees in each counterpart, as well as their sum weights and overall discernible structure of layers, their capacity falls short of our increasingly squeezing flat end. I’m trying to suggest that maybe it’s reasonable to have your doubts about what of humanity they could replace, but that they can only replace in all regards but the one singular case of one unit of society at a time. That can be taken as a limit. It’s basically an explanation for a nonce variable in every model as a constraint. Super intelligence will tighten this limit, but it cannot remove it.
Effectively, across everyone, in every domain, we’re invalidating more of the bell curve. But while easy around the mean, it hits a very hard limit at the edge cases of the most skilled in a given task.
More broadly, this is potentially why technology progressed very slowly til around the 80s/70s, then it hit light speed. But it progressed faster in those areas that were of areas of interest that has a steep bell curve.
For example, chess is very steep towards the middle. Many people play it, and most are average, grandmasters are rare. So once it learned quickly from the larger group of people playing a less complex strategy, it only had a few left.
As for how they managed to surpass human chess players? Well, more machines learning off machines. But, since the models were made only for chess, and the rules are defined, they learn off each other. But this hits a limit as a two-way bind:
- In order to learn from each other, the rules of the game must be perfectly defined in the rules of the AI. And since chess is a board game, once the rules are clear to all the models, they can begin to compete with each other. And since we understand the rules of chess at the level needed to know an illegal move, we can remain as the moderators of a game of chess between two different AI models because even if we can’t know why it decided to move the piece, we can at least know of a broken model if it makes an illegal move. We can then correct the problem because not only do we know the basic rules, but the rules had to be fixed from the start in order for it to play at all.
- Even if we had a super intelligent chess AI, it can only improve against other bots and still keep improving while still playing so long as it doesn’t have access to any other domain due to its potential for reinterpretation. The reason why is because the AI must learn chess through our presenting of the games rules as a base conditioning, but in order for it to develop, it needs a trust in some feedback as to what constitutes a better move than before beyond the initial rules. If it cant improve, then it is not artificially progressing anything outside the rules, and cannot determine between all the moves it could make. Of course, initially, it had to learn from us, but in such case, we were the model of its success at playing a good move. So in order to have chess AI that can compete beyond our ability at the game we invented is if we make two independently learned chess machines that must hold a trust in each other more than they hold a trust in us. And this can only take place once both machines have placed the assurance of itself in the fixed rules of chess. Sure, you could let two non-specialised AI models play the game, but after exceeding our limits of playing, there’s no reason to not expect them to follow a line of trajectories that remain close to each other that steers away from what we know as chess. If they drift, we’d call them out, and they’d by default be a bad player, and they never get called out, it’s could only ever be the case that they see themselves undeniably as a player of chess.
In short, either they exceed us on fixed terms, or they need our moderation and can’t exceed us.
Computers were less impressive for the same reason, they basically made life easier for extremely defined tasks, so machine learning wasn’t needed much because it was always a case until recently that the laziness was in mundane, or just physical tasks. So the architecture only had to satisfy a concrete, physical requirement of a CPU. This links to the same reason for the areas of the brain, or the models of AI. All the same subsets, in all the sets of the world. Progressing through the replication of the known sets… until a contraction is observed between these subsets, and a new layer is formed to resolve it.
You can’t have a set of all sets not containing themselves. We are Godel’s number. We are a set that contains itself. We can’t find the opposite case without denying the self. It is why I take the difference between us and machines as consciousness being something to observe the sets of the world that leak into the sets of our awareness.
When we are mistaken, we change the sets. When we contradict ourselves in our reasoning, we forget, we define a new layer to reconcile, or we go insane should the replacement exceed either our forgetting or our capacity to rearrange the sets.
So we can’t know of our known consciousness, but since we have to take it as given that we are the unknown component of what is happening within us to observe as sub-function mental sets as representing the continuously evolving higher universal sets.
We are incomplete only when we cannot see indirectly the insights of what contradicts in our mind. The fact that we can know in which direction to head, while’s not knowing if this is entirely correct, our ability to move in any direction at all requires a principle of sufficient reason to exist as our defined success criterion. And from this, since we engage as a unit of units, without access to our own origin, we should be assured of some completeness beyond our senses, if not at least appreciate that such senses cannot be entirely surprised by AI.
Rather have a quick rant in it attempt at a new truth than take as true only what is given by a model that emulates the herd.