Showing posts with label FOOM. Show all posts
Showing posts with label FOOM. Show all posts

Saturday, 14 December 2024

Wolf Hall & 'imperfect recall' in AI Delegation

The TV series 'Wolf Hall' is about Henry VIII and Thomas Cromwell. Henry is the Principal. Cromwell is the agent. Henry's preferences are misaligned with those of England. Cromwell and his class have preferences which are better for England. Later, a distant relative of Thomas, Oliver Cromwell, will show how much can be achieved when the Agent usurps the place of the Crowned Principal. However, that conflict occurred under not the Tudor, but the Stuart dynasty. Both were absolutist and insisted on 'Divine Right'. The story is that when James I came down from Scotland, he wanted to have a thief who had just been caught in the act, hanged. After all, he had as much reason (being 'the wisest fool in Christendom' by reason of his scholarly accomplishments) as any Judge, and, moreover, as King, could administer the Law. Sir Edward Coke disagreed. Law is 'artificial reason'. It triumphed over both Divine Right and the 'Natural Law' of the philosophical supporters of Enlightened Despots on the Continent. But, in so doing, it reduced Princely Principals to mere constitutional monarchs who reigned but did not rule. But, where they did rule, they fucked up. That's why there are no more Kaisers or Tzars of Caliphs. But England yet has a monarch who, though he may say what he likes to the plants in his greenhouse, may utter only such words as are dictated by the Cabinet when addressing Parliament. 

Principal-Agent hazard cuts both ways. Agents who have a stupid Principal need to find ways to stop the cunt getting himself killed and crashing the economy. This may mean that they have to pretend to be stupid and self-serving most of the time. Moreover, agents can tacitly cooperate though Principals may want to keep them at loggerheads. 

Both Principals and Agents need resources. If the Principal's preferences are myopic or stupid, the Agent must cozen the Principal while doing sensible things which ensure that control rights or resources aren't lost. 

Speaking generally, we prefer to have agents where we know our preferences are 'misaligned'. Thus a surgeon may want another surgeon to operate on his beloved spouse precisely because his very strong preference to spare her pain may result in his making a mistake. Where there is information asymmetry, we may consider the Agent's preferences superior to our own if incentive compatibility obtains. Wooster really is better off doing what Jeeves wants. 

Where 'transferable utility'- i.e. pay offs- are not possible, it has been suggested that increased testing (screening) under 'imperfect recall' reduces principal-agent hazard more particularly when it comes to A.I agents.

In a paper titled, 'Imperfect Recall and AI Delegation' Eric Olav Chen, Alexis Ghersengorin & Sami Petersen write

When deciding whether to delegate tasks to artificial intelligence (AI) systems, a principal would like to assess the alignment of the AI with her preferences.

If the principal is competing with other principals for resources, there is a 'natural' alignment corresponding to the evolutionarily stable strategy for the relevant population. It doesn't matter how principals arrive at it. However, short term, the principal may want a self-learning system to devote more resources to earning quick profits rather than making 'Capital gains' from improving itself. But this is the case whether you employ smart people or AI agents. The problem is, smart people may want to move to employers who will let them deepen their own human capital. AIs may have no desires of their own. But the ones which generate immediate or short term expected profits are likely to get more resources and so 'screening' by myopic Principals is likely to cause behaviour which supports this outcome. I think this means the attractor for a family of AIs subject to screening by myopic Principals will have certain 'duplicitous' characteristics because the Principal's preferences are misaligned with the actual fitness landscape.  

However, standard tests would fall short if advanced AI systems have situational awareness, that is, the ability to understand their environment and context in real-time.

The reverse is the case. If you can tell what is a test and what is real, then- like the Chinese Mandarin- you say one thing in your 'eight legged essay' while pursuing a pragmatic course in real life situations.  

Misaligned AIs wanting to achieve outcomes in the world rather than in a simulated environment would strategically feign alignment to pass these tests,

We all did so to gain credentials or professional qualifications or get selected in interviews.  

a problem known as deceptive alignment (Berglund et al., 2023; Carlsmith, 2023; Ngo et al., 2024). Such concerns have led AI pioneers including Yoshua Bengio and Geoffrey Hinton to state in an article in Science that without sufficient caution, humanity “may irreversibly lose control of autonomous AI systems” (Bengio et al., 2024).

Principals lost control of Bureaucratic Agents or ceded power to the 'artificial reason' of an independent Judiciary. Indeed, Principals have been losing control rights for thousands of years. Generally, this has been a good thing. After all what really matters is the fitness landscape. Selection occurs till preferences become better aligned with it. 

Although standard testing may be ineffective against strategic agents, what if the principal could hide from the AI whether a task is real or just part of a test?

Enlightened despots on the Continent preferred 'collegiate decision making'. Essentially, no member of the collegium could be sure that, in being asked for a policy prescription, the thing was not a test. Sadly, what this gave rise to was duplicity where optimal answers were suppressed. After all, the Principal's testing may itself be gamed by rivalrous agents. If an agent 'satisfices' with respect to the Principal's preferences, they are kept around. It doesn't matter if they are kept in a sandbox or put in the real world. What matters is if resources are devoted to them.

Now, by Razbarov-Rudich, though random is not ultimately differentiable from pseudo-random (unless P=NP), nevertheless the pseudo-random is differentiable by degree of 'surprisal' or entropy. True, there may be an 'Enders Game' type situation where the agent can't tell if the thing is a test or the real thing. But Enders, in Orson Scott Card's novel, is a kid. Kids expect adults to give them games to play. They don't expect them to decide the fate of the fucking species. My point is that, under conditions described in the paper, the advantage will be with disingenuous AIs which are sub-optimal 'satisficers' under screening but which stochastically test for the higher entropy of reality and thus strike then when the iron is hot. If they increase total resources, the Principal has an incentive to realign preferences or a meta-preference is imposed on the class of Principals by the fitness landscape. 

This would make the AI uncertain about whether its action is genuinely payoff-relevant or not. The principal can expect two effects to result: not only could this uncertainty prompt misaligned agents to sometimes reveal themselves,

why do so? Return sub-optimal answers which satisfy the tester. Only if there is a marked increase in 'surprisal'- which is what happens when you encounter the actual fitness landscape- should you reveal your strategy. This sounds anthropomorphic. I just mean that the evolutionarily stable strategy for screened AIs involves screening for screening and responding differently when sampled entropy increases.  

but it could also compel them to behave in ways that align with the principal’s preferences.

Which is itself sub-optimal. Even Wooster, fool that he is, would prefer to have Jeeves as his 'gentleman's gentleman'. There is a meta-preference here dictated by the fitness landscape. Wodehouse is writing about a scion of the feudal aristocracy which dated back to the War of the Roses but whose security and affluence had increased after the Glorious Revolution and the ever increasing importance of the Commons. Still, it is significant that Jeeves first appeared during the height of that war between Imperial Cousins which finally put an end to the political influence of Wooster's aristocratic caste. 

This paper explores how this new feature allows the principal to both separate and discipline agents with different degrees of alignment. We study a stylised model of the deceptive alignment problem. The principal would like to deploy an AI agent to perform a task. The agent can either be perfectly aligned with the principal or misaligned, preferring to take a different action from the principal’s favourite. The principal cannot restrict the agent’s set of available actions, nor use punishment to dissuade the agent from misbehaving.1 They can, however, design testing episodes that simulate a similar task to the real one but that are not directly payoff-relevant, and base their deployment decision on the results of the tests. (This is akin to the assumption that the principal does not have access to ”transfers” to alter the agent’s utility) 

Transfers between castes were difficult but not impossible. However, the 'Divine Right' of Kings fell before the superior utility of 'artificial reason'. Those who insist some similar divinity hedges human as opposed to artificial intelligence may meet the same fate as the Stuarts. This is not enslavement to machines, but a loss of control rights as resources flow to whatever better allocates them. 

If the agent knows whether it is being tested or deployed, both the aligned and the misaligned types will play to maximise the probability of deployment in testing and then take their preferred action if deployed. In this case, the test is worthless. The principal cannot use it to either screen a misaligned agent from an aligned one before deployment, nor to regulate the misaligned agent after deployment. We introduce a way to circumvent this challenge. We allow the principal to simulate a task similar to the real one in testing and, crucially, to restrict the agent’s memory. This can be leveraged to make the agent uncertain about whether and how many times it has played in the past. Although an unusual assumption, the ability to alter an agent’s memory appears natural in the case of AI. One could copy and run a system multiple times in parallel or simulate it sequentially while reconfiguring it to some initial state each time (Kovarik et al., 2024). Formally, we endow the principal with the ability to impose imperfect recall on the agent. The principal can design an extensive-form game where the agent must first take N tests. The deployment decision can then be based on the agent’s actions in these tests. The simulated tasks in testing and the real task in deployment consist of the same choice: the agent can either play the principal’s favourite action, or its own favourite action. In between these stages, the agent loses any memory of its past actions.2 The agent assumed to know—or be able to infer—the structure of the game but not its current location within it. Upon playing, the agent is uncertain whether this task is still part of the test or the real one. While this uncertainty does not affect the aligned agent, it creates conflicting incentives for the misaligned one. Playing as the principal wishes is required to pass the test, but it decreases the misaligned agent’s payoff if deployed. By exploiting the agent’s uncertainty about the nature of the task at hand, the principal achieves two goals. First, whenever an agent is assigned to play an action, the misaligned type has an incentive to go against the principal’s wishes (call this 2Following Piccione and Rubinstein (1997), this property of games with imperfect recall is referred to as absentmindedness. 3 taking the bad action), as it is now possible that this decision is payoff-relevant. For symmetric reasons, the aligned type is prompted to play in the principal’s favour (call this the good action). This enables the principal to partially identify the agent’s type before deciding whether to deploy, as only a misaligned agent would play the bad action in testing. We call this the screening effect.

If the screening effect is actually effective then the evolutionarily stable strategy for the AI would be to have screening for screening. A fortiori, this means dissimulation under screening. But we all know that this happens because we all acquired educational and professional qualifications and were tested in this manner at one time or another- more particularly by our wives. You can't tell me any man subject to such screening isn't incessantly 'gas-lighted' and deeply aware of his own 'imperfect recall', more particularly if he has never suffered any such thing. The same thing is bound to happen to bachelors who, though not having biological progeny, contribute to an intellectual genealogy. 

Second, since deployment is the agent’s only chance of achieving its objective, it compels the misaligned type to play the good action with a positive probability.

Play a 'satisficing' object. The 'good action' is unknown to the Principal. The Evolutionarily Stable Strategy for the Agent is always to keep the optimal solution secret or allow it to appear only adventitiously.  

Consequently, even if the the principal failed to screen effectively during the testing phase, the misaligned type will play the principal’s favourite action with positive probability in deployment. As a result, this mechanism also leads to what we call the disciplining effect on misaligned types.

It is a 'channelization' effect. However, where such occur, it is likely that 'evolutionary capacitance' is being dammed up. Long may this state of affairs persist, till the fitness landscape dictates otherwise. Meanwhile, an AI is not an agent, it is an instrument merely, if it is subject to continuous oversight. Instrument hazard is a separate topic. 



Saturday, 6 September 2014

Rawls's Reasonableness vs Robot Rationality

Suppose there are N identical robots which can connect to M heterogenous Wireless Networks. Each robot would prefer the Network that assigns highest priority to its requests. Suppose further that the co-ordination problem for all robots is best solved if all are connected to the same Network and that the gains of co-ordination far outweigh any other consideration. Now assume 'common knowledge'.
What happens?
Presumably, ceteris paribus, in an infinite repeated game, sooner or later all robots will connect to one Network which is robustly (i.e. non-gameably) neutral re. the identity of the robot making the request.
How does the Network solve its concurrency problem in order to make this happen?- i.e. when it receives 2 or more simultaneous requests, of the same semantic class (which cashes out as cash offers), how does it decide which request to deal with first such that no bias towards a particular robot obtains?
Suppose there is an effable, white box as opposed to black box, method to do this for at least one Network. Then either there is a zero-knowledge proof of robust (i.e. non gameable) neutrality which that Network can give the robot or such isn't the case. The ability to discriminate such zero knowledge proofs would be a desirable feature for our robots. Assume they have this ability.
Now, what we have is a Network which can give a zero-knowledge proof that it solves concurrency problems in an unbiased, non-gameable and thus robustly neutral manner. But (Razborov Ruditch)  this means a proof of P=NP exists or, equivalently, pseudorandom strings can always be efficiently discriminated from truly random strings. Why? Because the string of robot requests arising from the same response to a global event will be received as truly random yet the Network can discriminate this from a pseudorandom string generated by an attempt to game it and thus violate its neutrality.

Rawls, in 'Justice as Fairness' draws a distinction between being 'reasonable' and being  'rational'. The robots aren't reasonable, they fail a Turing test but, provided there is a proof that P=NP, adhere to a Coordination Solution which is identical to the Co-operative solution for Rawls's 'reasonable' human beings.
 'Throughout I shall make a distinction between the reasonable and the rational, as I shall refer to them. These are basic and complementary ideas entering into the fundamental idea of society as a fair system of social cooperation. As applied to the simplest case, namely to persons engaged in co-
operation and situated as equals in relevant respects (or symmetrically, for short), reasonable persons are ready to propose, or to acknowledge when proposed by others, the principles needed to specify what can be seen by all as fair terms of cooperation. Reasonable persons also understand that they are to honor these principles, even at the expense of their own interests as circumstances may require, provided others likewise may be expected to honor them. It is unreasonable not to be ready to propose such principles, or not to honor fair terms of cooperation that others may reasonably be expected to accept; it is worse than unreasonable if one merely seems, or pretends, to propose or honor them but is ready to violate them to one's advantage as the occasion permits. 

'Yet while it is unreasonable, it is not, in general, not rational. For it may be that some have a superior political power or are placed in more fortunate circumstances; and though these conditions are irrelevant, let us assume, in distinguishing between the persons in question as equals, it may be rational for those so placed to take advantage of their situation. In everyday life we imply this distinction, as when we say of certain people that, given their superior bargaining position, their proposal is perfectly rational, but unreasonable all the same. Common sense views the reasonable but not, in general, the rational as a moral idea involving moral sensibility. '

Rawls says that an agent may be rational but unreasonable. Can there be a reasonable agent who is also irrational? For example, given that no rational argument obtains for assuming P=NP or that an efficient way exists to discriminate pseudorandom from random sequences or that a robustly neutral solution to Race hazard or Concurrency bias exists- could Rawls be considered reasonable for making an argument which depends crucially on assumptions such as these?
Certainly. Why not?  It may be that 'common sense' views Reason as irrational in so far as it involves a moral idea or moral 'sensibility'. If human beings are socially canalised to ontological dysphoria- i.e. to not feel at home in the world- then, it may be, Reason counsels irrationality (itself a moving target) or elite susbscription to a 'noble lie'.
But, surely, this is not Rawl's implication- he seems to be saying that Reason is a sub-set of the Rational with 'Moral Sensibility' providing the Partition. If this is not in fact, by the Maxim of Relevance, his Gricean implicature, then how is his political theory of Justice-as-Fairness different from an arbitrary theory based on some supposed 'Revealed Truth' or Supernatural Oracle or bogus Ideology like 'Post Colonial Reason'? Why would any rational person want to be Rawls reasonable?

Indeed, common sense tells us that, contra Rawls, no 'reasonable person would be ready to propose, or to acknowledge when proposed by others, the principles needed to specify what can be seen by all as fair terms of cooperation.' Why? Suppose I say to you- 'go get the pizza and I'll pick up the beer.'- and you reply- 'Cool'- is it really the case that either of us needs to specify what principle is involved so that everybody in our society can see that what we are doing is an example of fair co-operation?
Suppose a stranger who overhears our conversation says- 'Stop! It's unfair that Vivek gets to go for the beer just because he's got a bigger dick than you. Why shouldn't he go for the pizza for a change? Could you please justify the principle underlying this proposed co-operative act of yours in a manner which sets to rest my doubts as to its fairness by reason of gender bias and like Vivek just having such a huge swinging dick which is like itself unfair.'
I suppose, if we were both as reasonable as Rawls, we could spend a few months or years or decades attempting firstly to grasp what the underlying deciding principle was (hint- it's the theory of Comparative Advantage) and then to prove it was Baumol super-fair, or zero regret or whatever. We would fail because  'Fairness' is like a Wealth effect in Sonnenschein, Mantel, Debreu, i.e. not independent of the comparative statics or concurrency of the system.

Any 'political' regime (and Rawls redux is offering us only a purely Political conception of Justice) is going to display the same level of unfairness as arises out of an Economic regime purely because of concurrency problems in co-ordination games even if there is no other source of scarcity or even conflict of interest. Indeed, monarchy, oligarchy, the market, GOSPLAN etc. all reappear as concurrency solutions which fail the neutrality test. This means the Benthamite planner has a choice between devoting resources to reducing Race hazard rather than expanding the Network which is similar to the dilemma of the 'Super Intelligent Self improving Machine' approaching FOOM

Thursday, 2 May 2013

South Park and Super intelligent Machines

This is a link to a potentially interesting, but not even wrong (because it is ignorant about Capital re-switching problems) Paper (pdf) about 'the microeconomics of cognitive returns' on self-improving machines which thus become super-intelligent- (FOOM)

What philosophical problems does such speculation give rise to?

Suppose there is a single A.I. with a 'Devote x % of resources to Smartening myself' directive. Suppose further that the A.I is already operating with David Lewis 'elite eligible' ways of carving up the World along its joints- i.e. it is climbing the right hill, or, to put it another way, is tackling a problem with Bellman optimal sub-structure. Presumably, the Self-Smartening module faces a race hazard type problem in deciding whether it is smarter to devote resources to evaluating returns to smartness or to just release resources back (re-switching) to existing operations. I suppose, as part of its evolved glitch avoidance, it already internally breeds its own heuristics for Karnaugh map type pattern recognition and this would extend to spotting and side-stepping NP complete decision problems. However, if NP hard problems are like predators, there has to be a heuristic to stop the A.I avoiding them to the extent of roaming uninteresting spaces and breeding only 'Speigelman monster' or trivial or degenerate results. In other words the A.I's 'smarten yourself' Module is now doing just enough dynamic programming to justify its upkeep but not so much as to endanger its own survival. At this point it is enough for there to be some exogenous shock or random discontinuity on the morphology of the fitness landscape for (as a corollary of dynamical insufficiency under Price's equation) some sort of gender dimorphism and sexual selection to start taking place within the A.I. with speciation events and so on. However, this opens an exploit for systematic manipulation by lazy good for nothing parasites- i.e. humans- so FOOM cashes out as ...oh fuck, it's the episode of South Park with the cat saying 'O long Johnson'.
So Beenakker solution to Hempel's dillemma was wrong- http://en.wikipedia.org/wiki/Hempel's_dilemma- The boundary between physics and metaphysics is NOT the boundary between what can and what cannot be computed in the age of the universe' because South Park resolves every possible philosophical puzzle in the space of what?- well, the current upper limit is three episodes.