Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5815_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
20 Мб
Скачать
Appendices 141
factor 1/2. Conversely, to find the hor i­zontal asymptote for unit 2 in the presence of unit 1, draw a vertical line from μ
to
1
intersect the distribution for unit 2, and scale by 1/2.
App 1.8.4 Bayesian Races
The problem may also be approached from a Bayesian perspective. Suppose we have two hypotheses H need not be mutually exclusive, and some evidence E; let the probability of getting result E on hypothesis H be p(E|H). Then the likelihood of H given E is L(H|E), = kp (E|H) where k is an arbitrary constant (Fisher 1921), which then cancels out. So the likelihood ratio is then L(H E); the log of this quantity has been called the weight of the evidence afforded by E (W (H pared with H
: E)) in favour of H1as com-
1/H2
(Good 1950, Good 1968,
2
Good 1975):
WðH
: EÞ¼log LðH1,H2jEÞ
1=H2
¼ log pð E jH
and H2, which
1
|E)/L(H2|
1
Þlog pðEjH,
1
(App 1.18)
with the advantage that the combined weight of two pieces of evidence is simply the sum of their individual weights.
App 1.8.5 Races between Many LATER Units
If we have a large set of possible hypoth­eses (for instance, the possible existence of several stimulus objects), we do not want to have to compare each with every other one in pairs. Rise-to-threshold provides a means of simultaneously comparing all hypotheses and selecting the most likely one, if each decision signal is taken as a measure of the belief B hypothesis H
|T)), and K is an arbitrary constant,
(L(H
j
, where Bjis equal to K + log
j
so that the log-odds log(p(H any pair of hypotheses is given by (B
B
). Figure App 1.25 represents a scheme
j
of this kind; it has a set of hypotheses H and a selection E composed of a selection from m possible stimulus elements S related by a matrix L
). Thus the log-likelihood for H
L(H
j|Si
given E is ΣI(log Lij), and the updated posterior belief is B
in an associated
j
)/(HLj)) of
j
of the likelihoods
ij
0
= Bj+ Σi(log Lij)
j
j
j
,
i
j
Figure App 1.25 Schematic representation of parallel multiple implementation of the rise-to-threshold model. A stimulus event E consists of the detection of a set of stimulus elements S set of potential hypotheses H weight of the evidence from E, is calculated from the learnt associations p(S hypotheses, and used to update the corresponding belief function B reaches a criterion level, at which point the appropriate response is initiated and all the Bjare reset.
concerning the existence of particular targets, of which there are n, the sum of the
j
, of which there are m. For each of a
i
) between stimulus elements and
i|Hj
. This process continues until one of the B
j
j
142 Appendices
If we are not interested in the odds for individual pairs of hypotheses, but only in the best overall hypothesis, then all we need do is run a race. B
starts at some level
j
representing the prior likelihood, and then, if updating occurs at constant intervals of time, E will rise at a rate proportional to the log of the likelihood ratio, until it reaches a criterion level which may be taken to repre­sent a sufficient degree of belief to permit action. The dynamic properties of the model can be taken care of simply by arran­ging for feedback such that the B
0
1 is equal to B
at time t.
j
at timet +
j
App 1.9 Learning
In Section 4.1 in Chapter 4, we saw that a change in prior prob ability in the middle of a run causes the reaction time on each side to alter in much the same way as a staticprior probability that is constant throughout the run. The time-course of this alteration appears to be roughly expo­nential, taking about 70 trials finally to settle down at the new equilibrium. It was pointed out there that the phenom­enon is better considered as a process of forgetting, in that the contribution of old evidence gradually contributes less new, being discounted by what we called a Lethean factor λ, whose value is typically around 0.05. This is not, of course, how a strict Bayesian model should behave: all evidence, whether recent or not, should count equally. It is, however, relatively easy to model at the synaptic level, using a Hebbian-like mechanism.
An alternative formulation (Peirce
1878) just needs two opposed neurons coding for I(H) and I(~H), an event E increasing the firing of each in propor­tion to I(E|H) and I(E|~H) respectively. Synapses must therefore strengthen in proportion to Peircean probability of E with respect to H and ~H.
The distinction between probability (for hypotheses) and chance(for events)
unnecessary: both are talking about the same kind of neural activity: to a neuron, so an event is a hypothesis. It is also clear that probability and causation are the same kind of phenomenon. And so is perception itself: the reconstruction of the realworld is as much guesswork as the extrapolation into future possible worlds, and involves precisely similar neural processes. Though we tend to com­pare the uncertainty of the future with the certainty of the past, as every historian knows, we guess the past just as much as the future.
Probability is represented in two quite different ways: strengths of synapses rep­resent conditional probabilities or likeli­hoods, while firing frequency represents degree of belief in a hypothesis. The pat­terns of activity of afferent synapses rep­resent patterns of circumstantial support, while sub-threshold depolarisation repre­sents prior probability represented. Then the Lethean factor λ (Chapter 4, Section
4.1) represents the speed with which syn­aptic strength can change.
App1.9.1 HebbianSynapsesas Bayesian Computers
We have seen that in an uncertain world decisions are all about probability, and showed how sensory inf ormation is used as evidence and determines the strength of ones belief in hypotheses about the out­side world, and how to quantify this process.
The final stage of this whole process is something called confirmation, which is at the heart of how neurons learn. So far, the various conditional probabilities p(E|H) have been presented as if they were given, but of course they have to be learnt through experience, being themselves updated when the truth about whether the prediction was actually correct is finally revealed. This is exactly equivalent to the use of parametric feedback after the
Appendices 143
Figure App 1.26 Highly simplified representation of Pavlovian conditioning. In the untrained animal there is an intrinsic, hard-wired link by which food (the unconditional stimulus, or UCS) causes salivation (a). After sufficient of food with a conditional stimulus (CS) such as a bell, the CS will trigger salivation even in the absence of food (b). The inescapable conclusion is that there are now two chains of neurons, from the CS and UCS, and that there must be at least one neuron (X) that is common to both (c). If this neuron shows Hebbian learning, its connection from the CS will be strengthened (d).
Figure App 1.27 (a) Chains of neurons connecting CS and UCs to the response R. (b) Neuron X has a Hebbian synapse ultimately driven by the CS. (c) If we identify U with a hypothesis H, and E with an observed event (the conditional stimulus CS), then the strength of the synapse (p(E|U) represents likelihood and therefore embodies Bayesian learning:
event in motor control and is quite an enlightening way at looking at many kinds of learning.
Take the simplest of all – Pavlovian conditioning (Pavlov 1927), Figure App 1.26. After the conditioning has been learnt, the dog is in effect using the bell as evidence E for the hypothesis U that food is in the offing. And when it does finally appear, the value of p(E|U), p(bell|food) is increased. This updating or confirmation is what psychologists mean by reinforce­ment, but it is obviously an example of a Bayesian process as well.
Learning of this kind can be per­formed by Hebbian synapses (Hebb
1949), of which NMDA synapses are a well-known example. When they get stronger as a result of association between pre-synaptic and post-synaptic activity, this is equivalent to altering p(E|H), so that next time the hypothesis is con­sidered even more likely when the bell is heard, Figure App 1.27. Things that tend
to happen together in the outside world tend to get associated together in the brain: fire together, wire togetherWe are building in our brain a model of prob­ability relationships in the outside world, perpetually predicting whats coming next (Carpenter and Williams 1995, Brodersen, Penny et al. 2008), Figure App 1.28.
Can we work out what the rule for Hebbian strengthening (and weakening) must be if a synapse is to function as a Bayesian element? If it acts linearly, its strength S should be proportional to log (p(U|C)). Its history consists of the number of instances of U andC (n
UCS
occurring together, and the total number of instances of C on its own (n S = log(n
UC/nC
).
). Then
C
So how must S change in response to C or UCS in order to generate this func­tion? The answer is that after every occur­rence of C the strength should decline by log(1 + 1/n
), and if U occurs as well it
C
)
144 Appendices
Figure App 1.28 Thanks to Hebbian synapses, the connections between neurons in the brain come to correspond to neural connections forming a model of the outside world.
should also increase by log (1 + 1/nUC). Since log (1 + x)=x – x
2
/2 + x3/3, ... ,to a first approximation the strength should change by –1/n
and 1/nUC, respectively.
C
Unfortunately, this implies that the syn­apse needs to have a memory not only of its own strength, but also of its tally of UC and C events. This may sound unrea­sonable: what we want is an updating rule that uses only the current S and the fact of the event U&C or C. However, the overall strength of the synapse might be the result of two parameters representing independently the history of C and the history of UC. They could be the numbers of two different kinds of mem­brane channel, for example, NMDA versus AMPA: is it plausible that the total excitation could be a log function of the number of active channels? If S =log
) – log(nC), then the rule could be
(n
UCS
extremely simple: after U&C, the number of each type of channel increases by one; after C alone, n
increases by one. So,
C
one might predict two sets of channels, one excitatory and increasing after con­junction, the other inhibitory and increasing after presynaptic activity only, though this is not in fact how AMPA and NMDA receptors behave. Physiologists will recogn ise that the formula for S is in effect the same as for the Nernst potential (V = k(log(C
), where C
1/C2
1
and C2are the numbers of each of the ions on each side of the membrane.
App 1.10 Information and Probability
App 1.10.1 Uncertainty as Lack of Information
Another conceivable scenario regarding the input of information regarding a hypothesis is that some information is provided about this hypothesis, but then a contradictory signal is given that cancels that previous message. Such conflicting input can be encoded as illustrated in Figure App 1.29.
App 1.10.2 Information in Extended Displays
In Section 4.6 in Chapter 4, we looked at various types of tasks in which the subject is required to make a judgement about whether in a field filled randomly with a mixture of two or more individual cat­egories of discrete stimuli (for instance, red and green dots), there are more of one of the categories than the other. In the Type 1 version of this task, a propor­tion a of the total of N items are the same, while the remainder are random (for example, in an RDK experiment a dots may move consistently to the right, while the others execute a random walk). In a Type 2 version, there are a items of one kind (for example, moving to the right), and the remainder (N – a) are of the other kind, so that discrimination gets more difficult as a approaches (N/2). A simple
Appendices 145
Figure App 1.29 (a) Initial probability for some hypothesis H. (b) A message is received telling us the value B of p(H). (c) A second message is received, saying that the previous message was untrue. P reverts to its original (red),
but the total path length is increased.
example of a Type 1 task was addressed in Reddi and Carpenter (2003). One needs to bear in mind that estimating the informa­tion content of such displays is not entirely straightforward and depends on certain assumptions, in particular on how local the estimation of direction of movement for any one of the detector units is. Here we assume it to be very local, as also have Weiss and Adelson (Weiss and Adelson 1998, Reddi, Asrress et al. 2003); the topic has been thoughtfully discussed by Barlow and Tripathy (Barlow and Tripathy 1997).
We start with the two opposed hypotheses, that at a particular moment the majority of dots are moving to the right (H
) or to the left (HL). Using
R
Bayes, the observation E of one dot moving to the right will increase the log likelihood ratio for H
versus HLby log
R
(C), where C is the likelihood ratio, (p(E|
)/(pE|HL).
H
R
In a Type 1 experiment, with N items, the probability of a particular item moving rightwards if H
is tr ue is (aN +
R
(N – aN)/2)/N,or(1+a)/2. Similarly, the
probability of an item moving leftwards if
is true is (1 – a)/2; so the likelihood
H
L
ratio for any one item will be C =(1+a)/ (1 – a): in many ways, C can be thought of as a kind of velocity contrast. The support for H
against HLfor the entire display
R
will be N log C, so in terms of LATER the median rate of rise of the decision signal be proportional to log (C). Then the median reaction time will be
þ k=ðN log CÞ,
T
0
where k is an arbitrary constant that is likely to vary from person to person, and
is the constant delay encapsulating all
T
0
those factors, such as conduction time, synaptic delay, and the time needed to activate muscle, that can be regarded as constant for any particular task. The larger a is, the shorter the reaction time will be.
In a Type 2 experiment, the difference is simply that the probability of observing a parti cular dot moving rightwards if H is true is a, but (1 – a)ifHLis true. So, C is now given by a/(1 – a).
R
Appendix 1 Mathematical
Mathematics is the art of giving the same name to different things. Henri Poincaré, Science et méthode (1908)
For those who appreciate such things, this Appendix brings together the more purely mathematical aspects of what has been discussed in the preceding chapters.
App 1.1 Notation
App 1.1.1 General
t Time s Reciprocal time, or promptness i trial index
Reciprocal latency: set of
s
i
promptnesses
N Number of trials
Latency, set of latencies
T, T
i
Median latency
T
M
L(t) Probability density function for
latency
R(s) Probability density function for
reciprocal latency
C(s) Cumulative frequency of s C(0) Terminal frequency of s P(x) Cumulative normal function Q(x) Complement of the cumulative
normal function: Q =(1 – P)
S Decision signal
Initial value of S
S
0
Threshold value of S that initiates a
S
T
response
θ Range of S:(S
r, r
Rate of rise of S, set of rates of rise
i
μ Median of r
2
σ
Variance of r
– S0)
T
i
i
δ Delay τ Time constant
E(x) Expected value of x
i
i
App 1.1.2 Specific Terminology for Inference
In addition to probability itself, there are many subsidiary concepts that need to be name and defined. The following does not claim to be exhaustive: for a more com­plete list, covering standard notation in information theory, see Good (1955). It is unfortunate that different authors have often introd uced new names for entities that already existed.
| given
: provided by
modified by observation
, list separator for
propositions
X, Y, A, B, etc. propositions
H hypothetical proposition
~H Negation of H
H
1,H2,H3,
E event
p, q probabilities
odds(A,B) p(A)/p(B)
lod(A,B) Log odds log (p(A)/p(B))
ent(H) expected
ent(p) total
ent(H:E
... Composite hypothesis
information concerning H
expected information
) entropy
j
concerning H provided by E
j
E I(H)
p log(p) – (1 – p)log (1 – p)
EjΣip(Hi|Ej)I (H
).
i:Ej
115
116 Appendices
In Shannon terms:
ent (H:E)
H(x) ent(H)
H(x,y) ent(H.E)
H
H
ev(H:E) Experimental
I(H:E) support log(p(E|H)/p(E))
I(H
cred(H) ev(H:E) + cred(H)
U(X) impossibility U(X) = 1/p (X)
Rate of transmission, R
(y) ent(E|H)
x
(x) ent(H|E)
y
Expectation of ent (H:Ej) before Ej observed
support
:E) composite
i
support
same as Jeffreysimprobability, imp(X) = 1/p(X) (Jeffreys 1931)
–log(p (X) identical with (Good 1955, Good 1971) Goods I(X)
p(Ej) ent
Σ
j
(H:E
) = ent
j
(E:H)
sur(E|~
log(p(E|Hi)/p(E)) = log(p(E|H
Σ
)/
i
(p(E|Hj)p(Hj))
j
support log
likelihood ratio (LLR)
log(p(E|H)/p(E| ~H)) or I(E|H) – I(E|~H)
support I(H:E) log (p (E|H)/p
(E))
Belief increases linearly in a series of
identical trials with identical outcomes (the support is constant). In a strict Peircean formulation, ~H is tricky, because there may be no corresponding p(E); but provided there is a set of actual alternative hypotheses to choose from, neural implementation is not a problem
So (I0– I) or log(p/p0) is a measure of information gained or lost through experience, per axis.
sex(p) (Weaver 1948)
surprise index
In logarithmic form, E (log p) – log p
E (~p)/p: (expected ~p divided by p actually observed.
amount of info in event, less the expected amount.
W(H1, H
:E)
2
weight of evidence concerning (H1v. H2) provided by E
)/p(E|H2) = I(H1:
1
:E)
2
(1/x)dx,=–log(p)
Information I =
log(p(E|H
E) – I(H
Ð
1
p
(workdone in moving from one probability to another, realised when found true).
bel bel as the only measure of
probability (the most fundamental)
likelihood p(E|H)
likelihood
p(E|H)/p(E|~H)
ratio
Rational support
I(E|A) – I(E|B)
- this is Edwards support
for A vs. B
Global support for A
Peirce support for A
I(E|A) – I(E)
I(E|A) – I(E|~A)
- same as corroboration, above
- same as confirmation, above, or weight of evidence
The ranges of the some of these func-
tions are:
p 0to1 U 1to
bel to +
Appendices 117
imp 1 to ent(0) ent(0) = ent(1) = 0; ent(0.5) = 0.5;
ent(0.25) = 0.5.
App 1.2 Properties of the Recinormal Distribution
In general, a normal distribution is described by the frequency or probability density function, P(x):
2
xμðÞ
1
2
1
ffiffiffiffiffiffiffiffi
2πσ
2σ
, (App 1.1)
e
x
ð
2
zμðÞ
2
2σ
e
dz: (App 1.2)
ffiffiffiffiffiffiffiffi
PxðÞ¼
where μ is the mean of the distribution and σ
p
2πσ
2
is the variance (σ is then the stand­ard deviation, a measure of the overall width); μ is also the median value of x, since the distribution is symmetrical: P(x μ) P(μ – x) for all x.
Corresponding to this probability density function or PDF is the cumulative distribution function or CDF, which is given by:
CxðÞ¼
p
This is S-shaped (Figure App 1.1) and is normalised in the sense that it ranges from 0 to 1.
In a recinormal distribution, the recip­rocal of the variate is normally distrib­uted. In the case of reaction times, this means that the probability of a response in a particular trial having a latency T whose reciprocal (S =1/T) lies between s and s +ds is:
ffiffiffiffiffiffi
2π
sμðÞ
2
2σ
ds, (App 1.3)
e
2
is the variance.
1
RsðÞ¼
p
σ
where the mean (and also median) of the distribution is μ and σ
2
From this the probability density of the set of original reaction times we can derive T
LtðÞ¼
as
i
2
ffiffiffiffiffiffi
2π
1μtðÞ
2
2t
dt: (App 1.4)
e
1
p
2
t
σ
This is a positively skewed distribution whose median value T
,is1/μ, but whose
M
mean and variance do not have simple analytical forms.
Typical plots of R(s) and L(t) are shown in Figure App 1.2. When plotting reciprocal latencies, it is helpful to have the origin (s = 0) on the right rather than the left, so that latencies still increase to the right. It also aids comprehension to use a non-linear (reciprocal) scale of
Figure App 1.1 C(X), plotted cumulatively. Z is in units of standard deviation, σ.
118 Appendices
Figure App 1.2 (a,b) Frequency histograms of a simulated recinormal distribution (N = 5000, μ =5,σ = 1), plotting the probability density for the rawlatency T (a), and (b), for its reciprocal, S. Note that latency increases to the right in each case. (c,d) Cumulative frequency plots of the same distribution, using a linear ordinate (c), and a probit scale (d: a reciprobit plot), that generates a straight line if the distribution is indeed Gaussian. The line shown was fitted by minimising the Kolmogorov–Smirnov one-sample statistic.
latencies rather than a linear scale of reciprocal latencies.
If the recinormal distribution is
obeyed, a straight line will be obtained: if its slope is μ and it intersects the t = axis at an ordinate k, then μ = k/μ and σ =1/μ. Such a plot may be called a reciprobit plot, and is illustrated in Figure App 1.2(d). Probabilities are marked on the left-hand ordinate, probits on the right; the reci pro­cal time scale lies along the bottom, from right to left so that time increases to the right. A scale of reciprocal latency or promptness can be added above. Two crit­ical points on the line have well-defined meanings: where it passes through the horizontal 50% line is the median latency,
, while the intercept k represents the
Τ
Μ
probability that no response to the stimu­lus ever occurs at all. For easily detected stimuli k is typically of the order of 6–9, and the corresponding probability is so small as to be meaningless. But if the detectability of the stimulus is reduced to near threshold, for instance by reducing the contrast if it is visual, k will become a measurable percentage. k is also the ratio μ/σ, so that under a change in time-scale the line simply swivels around the intercept point.
For a normal distribution the standard
error of the median is 1.25 σ/N (Kendall and Stuart 1968); applying this to the recinormal distribution, the standard
Appendices 119
error of the median latency will be given approximately by multiplying it by 1.25/ kN. Thus, in a typical experiment, with a median latency of 200 ms, k = 7, and N = 400, the standard error is about 1.7 ms; such a figure can, however, be misleading, since measurements of successive medians of this kind shows them to be more widely scattered than Figure App
1.2 implies. Latency measurements seem to be subject to random fluctuations over long periods of time, with a noise power spectrum that shows some of the charac­teristics of 1/f noise (Gilden, Thornton et al. 1995, Wagenmakers, Farrell et al.
2004); for this reason, it is necessary to pay careful attention to experimental design when trying to compare the effects of different conditions, and to be aware that the standard error calculated for any one run underestimates the variation to be expected between runs.
App 1.3 Models for Latency Distributions
From time to time, but particularly in the 1960s and 1970s, attempts have been made to discover theoretical functions that could be used to describe reaction time distributions. Some have been derived from theoretical considerations, others purely empirically; some have many parameters, others admirably few. The review that follows does not claim to be comprehensive, or to attempt to com­pare their success in fitting observed data, or to assess their biological plausibility, but simply to provide a gallery illustrating their general characteristics. By plotting the functions as reciprobit plots, it is rela­tively easy to see in what ways they differ significantly from LATER-like behaviour. The most thorough way to do this would be simply to repeat the fitting of all the data that have been used here with the rival models, and see whether they per­form better or worse. But this would have
entailed a very great deal of labour, of a rather negative kind. The procedure actu­ally adopted was to take each of the alter­native models in turn, and display their distributions as reciprobit plots for differ­ent combinations of the parameters. In this way, one can establish combinations of their parameters that give the nearest approximations to particular recinormal distributions, and then run simulations with numbers of trials of the same order of magnitude as are typically used and see whether the results produce statistically acceptable fits. In this way, one can estab­lish which models are at least compatible with the recinormal distribution under ordinary conditions. Obviously, unless a model is analytically identical with a recinormal distribution, as the number of trials increases a point must eventually come where they are statistically incompatible.
App1.3.1 CountingModels
A general class of model that is intuitively attractive, and can sometimes generate quite realistic distributions, involves counting events that are generated sto­chastically and that generate a response when the counts reach some kind of cri­terion (Pike 1973). They give rise to dis­tributions that are closely related to the gamma and beta functions of stochastic theory, and it may be helpful first to sum­marise the underlying mathematics, with its rather messy notation: The (complete) gamma function is
ð
α1ex
Γ αðÞ¼
x
0
The corresponding incomplete gamma function has two parameters, and is given by
ð
γα, tðÞ¼
dx: (App 1.5)
t
α1ex
x
0
dx: (App 1.6)