F.Shannon — Channel Capacity as a Bound on Decision Quality¶
Area 01 · Decision-Making Under Uncertainty — F-Series Foundational Experiments
Historical thesis tested: Shannon's channel capacity theorem: a noisy binary channel supports reliable decisions at a rate determined by C = 1 − H_b(1 − p_correct). The capacity governs how many observations an agent needs, and how fast its uncertainty collapses.
This notebook asks: How does mutual information I(O;S) per observation predict decision quality? How many observations n* does an agent need to reach accuracy τ? What happens when the agent's model of its own sensor is wrong?
Key quantities (binary symmetric channel, BSC):
- Crossover probability: ε = 1 − p_correct
- Channel capacity: C = 1 − H_b(ε) [bits per observation]
- Mutual information: I(O;S) = C (at uniform input distribution)
- Log-likelihood ratio per obs: Λ = log(p/(1−p)) = logit(p) [nats]
- Exact accuracy after n i.i.d. obs: P(correct|n) = 1 − Binomial.CDF(⌊(n−1)/2⌋, n, p)
Experiments:
- E01: Exact theoretical formula (binomial) vs empirical accuracy
- E02: Cumulative log-LLR as a predictor of accuracy — approximate collapse
- E03: n* to reach accuracy τ vs channel quality
- E04: Belief entropy decay — how fast does uncertainty collapse?
- E05: Misspecified channel in a dynamic world — overconfidence causes tracking failure
Imports and configuration¶
import numpy as np
import matplotlib.pyplot as plt
from scipy import stats
from typing import List, Dict
np.random.seed(42)
COLORS = {
'p095': '#1e3a8a',
'p090': '#2563eb',
'p085': '#0891b2',
'p080': '#16a34a',
'p075': '#d97706',
'p070': '#dc2626',
'p060': '#9ca3af',
'theory': '#000000',
}
P_VALS = [0.95, 0.90, 0.85, 0.80, 0.75, 0.70, 0.60]
COLOR_LIST = [COLORS[f'p0{int(p*100):02d}'] for p in P_VALS]
plt.rcParams.update({
'figure.dpi': 120, 'font.size': 10,
'axes.spines.top': False, 'axes.spines.right': False,
'axes.grid': True, 'grid.alpha': 0.3,
})
Information-theoretic functions¶
def binary_entropy(p: float) -> float:
"""H_b(p) = −p·log₂(p) − (1−p)·log₂(1−p) [bits]."""
if p <= 0 or p >= 1:
return 0.0
return -p * np.log2(p) - (1 - p) * np.log2(1 - p)
def channel_capacity(p: float) -> float:
"""C = 1 − H_b(1 − p) for BSC with crossover prob 1−p [bits]."""
return 1.0 - binary_entropy(1.0 - p)
def theoretical_accuracy(p: float, n: int) -> float:
"""P(correct | n i.i.d. obs, uniform prior, BSC(p), tie-break → act=0).
Derivation:
Let X = #{obs = 0}. Under state=0: X ~ Binom(n, p). Agent acts 0 iff X ≥ n−X,
i.e. X ≥ n/2. Under state=1: X ~ Binom(n, 1−p). Agent is WRONG on a tie
(X = n/2) because tie-breaking always picks act=0.
By BSC symmetry P(X > n/2 | state=0) = P(X < n/2 | state=1).
P(correct | state=0) = P(X > n/2) + P(X = n/2) [ties → act=0 = state]
P(correct | state=1) = P(X < n/2) [ties → act=0 ≠ state]
Averaging over uniform prior (each state equally likely):
P(correct) = [P(X > n/2) + P(X = n/2) + P(X < n/2)] / 2
— but we need to be careful. With X ~ Binom(n, p):
P(correct|s=0) = P(X ≥ ⌈n/2⌉) = binom.sf((n-1)//2, n, p)
P(correct|s=1) = P(X ≤ (n-2)//2) = binom.cdf((n-2)//2, n, 1-p)
= binom.sf(n//2, n, p) [by symmetry of Binom(n,p) vs Binom(n,1-p)]
P(correct) = [binom.sf((n-1)//2, n, p) + binom.sf(n//2, n, p)] / 2
Key property: P(correct|2k) = P(correct|2k−1) for all k ≥ 1.
Adding an even-numbered observation NEVER improves expected accuracy.
The tie-breaking bonus (helps state=0) and penalty (hurts state=1) cancel exactly.
Effective sample size for binary decisions: ⌈n/2⌉ odd steps.
"""
return 0.5 * (float(stats.binom.sf((n - 1) // 2, n, p))
+ float(stats.binom.sf(n // 2, n, p)))
def theoretical_entropy(p: float, n: int) -> float:
"""E[H(b_n)] = Σ_{k=0}^{n} Binom(k; n, p) · H(σ((2k−n)·logit(p))) [bits].
After n i.i.d. BSC(p) observations starting from a uniform prior, the belief
b_n depends only on the sufficient statistic k = #{obs = 0} via:
b_n = σ((2k − n) · logit(p)) where σ is the logistic function.
Averaging over all possible k under state=0 (BSC symmetry makes state=1 give the
same distribution of H values) yields the expected entropy.
Key property: E[H(b_n)] decreases MONOTONICALLY with n — unlike decision accuracy,
which is pairwise-constant. Even observations DO tighten the posterior (reducing
entropy) even though they don't shift the majority vote.
"""
llr = float(np.log(p / (1.0 - p)))
total = 0.0
for k in range(n + 1):
prob = float(stats.binom.pmf(k, n, p))
L = (2 * k - n) * llr
bel = 1.0 / (1.0 + np.exp(-L))
total += prob * binary_entropy(bel)
return total
def log_llr(p: float) -> float:
"""Λ = log(p/(1−p)) = logit(p) [nats per observation]."""
return np.log(p / (1.0 - p))
Simulation: environment and agents¶
class StaticTwoDoorEnv:
"""Two-door environment with FIXED hidden state (p_switch=0)."""
def __init__(self, p_correct: float):
self.p_correct = p_correct
self.true_state: int = 0
def reset(self):
self.true_state = np.random.randint(2)
def observe(self) -> int:
return self.true_state if np.random.random() < self.p_correct \
else 1 - self.true_state
class DynamicTwoDoorEnv:
"""Two-door environment with switching hidden state."""
def __init__(self, p_correct: float, p_switch: float):
self.p_correct = p_correct
self.p_switch = p_switch
self.true_state: int = 0
def reset(self):
self.true_state = np.random.randint(2)
return self._observe()
def _observe(self) -> int:
return self.true_state if np.random.random() < self.p_correct \
else 1 - self.true_state
def step(self):
if np.random.random() < self.p_switch:
self.true_state = 1 - self.true_state
return self._observe(), self.true_state
class BayesianAgent:
"""Correctly specified Bayesian agent for static world."""
def __init__(self, p_correct: float):
self.p = p_correct
self.belief = 0.5
def reset(self):
self.belief = 0.5
def update(self, obs: int):
l0 = self.p if obs == 0 else (1 - self.p)
l1 = (1 - self.p) if obs == 0 else self.p
denom = l0 * self.belief + l1 * (1 - self.belief)
if denom > 1e-12:
self.belief = l0 * self.belief / denom
def act(self) -> int:
return 0 if self.belief >= 0.5 else 1
def entropy(self) -> float:
return binary_entropy(self.belief)
class DynamicBayesianAgent:
"""Bayesian agent for dynamic world: uses p_assumed for updates."""
def __init__(self, p_assumed: float, p_switch_assumed: float):
self.p_assumed = p_assumed
self.p_switch = p_switch_assumed
self.belief = 0.5
def reset(self):
self.belief = 0.5
def update(self, obs: int):
# Prediction step
b = self.belief * (1 - self.p_switch) + (1 - self.belief) * self.p_switch
# Update with assumed channel quality
pa = self.p_assumed
l0 = pa if obs == 0 else (1 - pa)
l1 = (1 - pa) if obs == 0 else pa
denom = l0 * b + l1 * (1 - b)
if denom > 1e-12:
self.belief = l0 * b / denom
def act(self) -> int:
return 0 if self.belief >= 0.5 else 1
def run_static_accuracy(env: StaticTwoDoorEnv, agent: BayesianAgent,
max_obs: int, n_episodes: int) -> np.ndarray:
"""P(correct at step t) for t = 1 … max_obs (agent accumulates obs within episode)."""
correct_counts = np.zeros(max_obs)
for _ in range(n_episodes):
env.reset()
agent.reset()
for t in range(max_obs):
obs = env.observe()
agent.update(obs)
correct_counts[t] += int(agent.act() == env.true_state)
return correct_counts / n_episodes
def run_dynamic_accuracy(env: DynamicTwoDoorEnv, agent: DynamicBayesianAgent,
T: int, n_episodes: int) -> float:
"""Mean accuracy over T steps in a dynamic world."""
total_correct = 0
for _ in range(n_episodes):
obs = env.reset()
agent.reset()
agent.update(obs)
total_correct += int(agent.act() == env.true_state)
for _ in range(T - 1):
obs, true_state = env.step()
agent.update(obs)
total_correct += int(agent.act() == true_state)
return total_correct / (n_episodes * T)
E01 — Exact theoretical formula vs empirical accuracy¶
Hypothesis: The exact formula P(correct|n) = 1 − Binomial.CDF(⌊(n−1)/2⌋, n, p) matches empirical simulation to within the noise floor 1/√n_episodes.
Key point: The Bayesian update with uniform prior is equivalent to majority vote: the agent acts on whichever state received more observations. Tie goes to state 0 (belief=0.5 satisfies the ≥0.5 condition).
Setup: Static world, n_obs ∈ [1,50], n_episodes=5000.
print("E01: Exact formula vs empirical accuracy")
print("=" * 50)
E01_MAX_OBS = 50
E01_N_EPISODES = 5000
n_arr = np.arange(1, E01_MAX_OBS + 1)
noise_floor = 1.0 / np.sqrt(E01_N_EPISODES)
results_e01_emp = {}
results_e01_thy = {}
for p in P_VALS:
env = StaticTwoDoorEnv(p)
agent = BayesianAgent(p)
np.random.seed(42)
emp = run_static_accuracy(env, agent, E01_MAX_OBS, E01_N_EPISODES)
thy = np.array([theoretical_accuracy(p, int(n)) for n in n_arr])
results_e01_emp[p] = emp
results_e01_thy[p] = thy
max_gap = np.max(np.abs(emp - thy))
within = "✓" if max_gap < 3 * noise_floor else "✗"
print(f" p={p:.2f} C={channel_capacity(p):.4f}b "
f"max|emp−theory|={max_gap:.4f} {within} (<3σ={3*noise_floor:.4f})")
fig, axes = plt.subplots(1, 2, figsize=(12, 4.5))
fig.suptitle("E01 — Bayesian Majority-Vote Accuracy: Theory vs Simulation", y=1.01)
ax = axes[0]
for p, color in zip(P_VALS, COLOR_LIST):
ax.plot(n_arr, results_e01_emp[p], color=color, lw=1.5, alpha=0.55)
ax.plot(n_arr, results_e01_thy[p], color=color, lw=2, ls='--',
label=f'p={p:.2f} C={channel_capacity(p):.3f}b')
ax.axhline(0.99, color='black', lw=0.8, ls=':')
ax.text(E01_MAX_OBS-2, 0.993, '0.99', fontsize=7, ha='right')
ax.set_xlabel("Observations n")
ax.set_ylabel("P(action = true state)")
ax.set_title("Accuracy vs n\n(solid=empirical, dashed=theory)")
ax.set_ylim(0.47, 1.02)
ax.legend(fontsize=7, loc='lower right')
ax = axes[1]
for p, color in zip(P_VALS, COLOR_LIST):
residual = results_e01_emp[p] - results_e01_thy[p]
ax.plot(n_arr, residual, color=color, lw=1.5, label=f'p={p:.2f}')
ax.axhline(0, color='black', lw=1)
for sign in [+1, -1]:
ax.axhline(sign * noise_floor, color='black', lw=0.8, ls=':', alpha=0.5)
ax.text(2, noise_floor + 0.001, '±1/√N (noise floor)', fontsize=7, color='#6b7280')
ax.set_xlabel("n")
ax.set_ylabel("Empirical − Theory")
ax.set_title("Residuals (should lie within noise floor)")
ax.legend(fontsize=7)
plt.tight_layout()
plt.savefig('e01_shannon_accuracy_curves.png', bbox_inches='tight')
plt.close()
print(f"\n Noise floor (1/√N) = {noise_floor:.4f}")
print("Saved: e01_shannon_accuracy_curves.png")
FINDING 1: The exact formula P(correct|n) = 1 − Binomial.CDF(⌊(n−1)/2⌋, n, p) is confirmed to within the simulation noise floor (1/√n_episodes ≈ 0.014) across all p values. Residuals show no systematic bias.
The majority-vote interpretation is key: the Bayesian posterior with a uniform prior integrates all observations as independent votes, and the decision rule (act=0 if belief ≥ 0.5) is equivalent to choosing the state that received more observations. A higher channel quality (larger p) shifts the binomial distribution so that the correct state more reliably wins the majority, at a rate determined by logit(p).
E02 — Cumulative log-LLR as a predictor of accuracy¶
Hypothesis: When re-indexed by cumulative log-likelihood ratio x = n · logit(p), the accuracy curves from different p values should approximately collapse. The exact collapse variable is the z-score z = √n · (2p−1)/√(p(1−p)), and P(correct) ≈ Φ(z/2) by the Central Limit Theorem for large n.
Setup: Use theoretical accuracy curves, compare collapse under two indexing variables.
print("\nE02: Cumulative log-LLR as predictor of accuracy")
print("=" * 50)
for p in P_VALS:
llr_per_obs = log_llr(p) # logit(p) nats per obs
cap = channel_capacity(p)
z_at_n50 = np.sqrt(50) * (2*p - 1) / np.sqrt(p * (1 - p))
print(f" p={p:.2f} logit(p)={llr_per_obs:.3f} C={cap:.4f}b "
f"z(n=50)={z_at_n50:.2f}")
fig, axes = plt.subplots(1, 2, figsize=(12, 4.5))
fig.suptitle("E02 — Cumulative Log-LLR as Predictor of Accuracy", y=1.01)
# Left: index by n × logit(p) (approximate collapse)
ax = axes[0]
x_max = 5.0
for p, color in zip(P_VALS, COLOR_LIST):
llr = log_llr(p)
thy = results_e01_thy[p]
x = n_arr * llr
mask= x <= x_max
ax.plot(x[mask], thy[mask], color=color, lw=2, label=f'p={p:.2f}')
# Approximate collapse curve: CLT approximation
x_plot = np.linspace(0.001, x_max, 300)
# P(correct) ≈ Φ(√n·(2p-1)/√(p(1-p))/2)
# In terms of x = n·logit(p): √n = √(x/logit(p)), so
# z/2 = √(x/logit(p)) · (2p-1) / (2√(p(1-p))) -- this still depends on p, not universal
# Better: exact for n=1 is p=σ(logit(p)), for large n it's approximate
# Show the identity line (n=1 point where all start) and just plot the collapse
ax.set_xlabel("Information index n·logit(p) [nats]\n(≠ realized LLR; E[LLR_n] = n·(2p-1)·logit(p))")
ax.set_ylabel("P(correct)")
ax.set_title("Approximate collapse by information index n·logit(p)\n(not exact — factor (2p-1) differs by p; CLT z-score collapses better)")
ax.set_xlim(0, x_max)
ax.set_ylim(0.48, 1.02)
ax.legend(fontsize=7)
# Right: index by CLT z-score (better collapse for large n)
ax = axes[1]
from scipy.special import ndtr # standard normal CDF
z_max = 5.0
for p, color in zip(P_VALS, COLOR_LIST):
z_n = np.sqrt(n_arr) * (2*p - 1) / np.sqrt(p * (1 - p))
thy = results_e01_thy[p]
mask= z_n <= z_max
ax.plot(z_n[mask], thy[mask], color=color, lw=2, label=f'p={p:.2f}')
# CLT prediction: Φ(z/2)
z_line = np.linspace(0, z_max, 200)
ax.plot(z_line, ndtr(z_line / 2), 'k--', lw=2, label='Φ(z/2) [CLT]')
ax.set_xlabel("CLT z-score √n·(2p−1)/√(p(1−p))")
ax.set_ylabel("P(correct)")
ax.set_title("Collapse on CLT z-score\n(better at large n)")
ax.set_xlim(0, z_max)
ax.set_ylim(0.48, 1.02)
ax.legend(fontsize=7)
plt.tight_layout()
plt.savefig('e02_cumulative_information.png', bbox_inches='tight')
plt.close()
print("Saved: e02_cumulative_information.png")
FINDING 2: The index n·logit(p) produces an approximate but imperfect collapse of accuracy curves across p values. Note: n·logit(p) is NOT the realized cumulative log-LLR. The realized cumulative LLR after n observations with k correct is (2k−n)·logit(p); its expectation is n·(2p−1)·logit(p), which differs from n·logit(p) by factor (2p−1). n·logit(p) is the LLR one would observe if ALL n observations were correct — an upper bound, not the expected value. It is used here only as an approximate indexing variable.
The CLT z-score z = √n·(2p−1)/√(p(1−p)) gives a better collapse because it accounts for the actual signal-to-noise ratio: P(correct|n) ≈ Φ(z/2), where Φ is the standard normal CDF. This approximation improves as n grows and degrades near p=0.5.
The difference between logit(p) and the CLT rate reflects a fundamental distinction:
- logit(p) = exact log-likelihood ratio per observation (Bayesian update rate)
- C = 1−H_b(1−p) = Shannon capacity (information-theoretic rate)
- CLT z-scale = (2p−1)/√(p(1−p)) (depends on both mean and variance of log-LLR) These three are monotonically related but not proportional. For p near 1 they agree; for p near 0.5 they differ substantially.
E03 — Observations needed to reach target accuracy¶
Hypothesis: The number of observations n* required to reach accuracy τ decreases with channel quality. From the CLT approximation: n*(τ, p) ≈ (Φ⁻¹(τ))² · p(1−p) / (p−0.5)²
Setup: Compute n* exactly (from exact formula) and via CLT approximation, for τ ∈ {0.90, 0.95, 0.99} over p ∈ [0.51, 0.99].
print("\nE03: Observations needed to reach target accuracy")
print("=" * 50)
E03_P_VALS = np.linspace(0.51, 0.99, 40)
E03_TAU_VALS = [0.90, 0.95, 0.99]
MAX_N_SEARCH = 500
tau_colors = ['#2563eb', '#16a34a', '#dc2626']
results_e03_exact = {tau: [] for tau in E03_TAU_VALS}
results_e03_clt = {tau: [] for tau in E03_TAU_VALS}
from scipy.special import ndtri # inverse normal CDF
for p in E03_P_VALS:
for tau in E03_TAU_VALS:
# Exact n*: smallest n such that theoretical_accuracy(p, n) >= tau
n_exact = None
for n in range(1, MAX_N_SEARCH + 1):
if theoretical_accuracy(p, n) >= tau:
n_exact = n
break
results_e03_exact[tau].append(n_exact if n_exact else MAX_N_SEARCH)
# CLT approximation: n* ≈ (Φ⁻¹(τ))² · p(1-p) / (p-0.5)²
z_tau = ndtri(tau) # Φ⁻¹(τ)
n_clt = (z_tau ** 2) * p * (1 - p) / max((p - 0.5) ** 2, 1e-10)
results_e03_clt[tau].append(n_clt)
# Spot-check table
print(f" {'p':>6} {'C':>8} {'n*(0.90)':>10} {'CLT':>8} {'n*(0.99)':>10} {'CLT':>8}")
for p in [0.55, 0.65, 0.75, 0.85, 0.95]:
idx = int((p - 0.51) / (0.99 - 0.51) * (len(E03_P_VALS) - 1))
n90_e = results_e03_exact[0.90][idx]
n90_c = results_e03_clt[0.90][idx]
n99_e = results_e03_exact[0.99][idx]
n99_c = results_e03_clt[0.99][idx]
print(f" {p:>6.2f} {channel_capacity(p):>8.4f} {n90_e:>10.0f} "
f"{n90_c:>8.1f} {n99_e:>10.0f} {n99_c:>8.1f}")
fig, axes = plt.subplots(1, 2, figsize=(12, 4.5))
fig.suptitle("E03 — Observations Needed vs Channel Quality", y=1.01)
ax = axes[0]
for tau, color in zip(E03_TAU_VALS, tau_colors):
exact_ns = np.array(results_e03_exact[tau], dtype=float)
clt_ns = np.array(results_e03_clt[tau])
ax.plot(E03_P_VALS, exact_ns, color=color, lw=2, label=f'τ={tau:.2f} (exact)')
ax.plot(E03_P_VALS, clt_ns, color=color, lw=1.2, ls='--', alpha=0.7,
label=f'τ={tau:.2f} (CLT)')
ax.set_xlabel("p_correct")
ax.set_ylabel("n* observations needed")
ax.set_title("n* vs channel quality")
ax.set_ylim(0, 150)
ax.legend(fontsize=7)
ax = axes[1]
caps = np.array([channel_capacity(p) for p in E03_P_VALS])
for tau, color in zip(E03_TAU_VALS, tau_colors):
exact_ns = np.array(results_e03_exact[tau], dtype=float)
mask = exact_ns < MAX_N_SEARCH
ax.plot(caps[mask], exact_ns[mask], color=color, lw=2, label=f'τ={tau:.2f}')
ax.set_xlabel("Channel capacity C [bits]")
ax.set_ylabel("n* observations needed")
ax.set_title("n* vs channel capacity C\n(n* ∝ 1/C near p=0.5 where C→0; CLT regime)")
ax.legend(fontsize=9)
plt.tight_layout()
plt.savefig('e03_observations_needed.png', bbox_inches='tight')
plt.close()
print("\nSaved: e03_observations_needed.png")
FINDING 3: The CLT approximation n* ≈ (Φ⁻¹(τ))² · p(1−p) / (p−0.5)² matches the exact formula well for p ≳ 0.65, diverging only near p=0.5 (where CLT converges slowly). n* decreases sharply with channel quality: doubling p_correct from 0.60 to 0.80 typically reduces n* by 3–5× for most thresholds.
The relationship n* ∝ 1/C holds in the SMALL-C regime (C→0, p near 0.5), where the CLT approximation gives (p−0.5)² ∝ C and n* ≈ const/C. This regime is NOT "large C": at large C (p near 1), n* falls quickly toward 1 but the 1/C proportionality breaks down because p(1−p)→0 and the CLT approximation itself fails for small n. What we can conclude: higher-capacity channels require fewer observations; the 1/C scaling describes the low-capacity limit near p=0.5, not the high-quality regime.
E04 — Belief entropy vs decision accuracy: continuous vs discrete¶
Hypothesis: While decision accuracy is pairwise-constant (staircase at even n), belief entropy E[H(b_n)] decreases MONOTONICALLY — each observation tightens the posterior even when it does not shift the majority vote. The exact theoretical entropy is E[H(b_n)] = Σ_k Binom(k;n,p)·H(σ((2k−n)·logit(p))). This differs from H_b(P(correct|n)) for all n ≥ 2.
Setup: Track mean H(b_t) per step across 3000 episodes, static world.
print("\nE04: Belief entropy vs decision accuracy")
print("=" * 50)
E04_P_VALS = [0.60, 0.70, 0.80, 0.90, 0.95]
E04_MAX_OBS = 40
E04_N_EP = 3000
E04_COLORS = [COLORS['p060'], COLORS['p070'], COLORS['p080'],
COLORS['p090'], COLORS['p095']]
results_e04_emp = {}
results_e04_thy = {}
for p in E04_P_VALS:
env = StaticTwoDoorEnv(p)
agent = BayesianAgent(p)
entropy_sums = np.zeros(E04_MAX_OBS)
np.random.seed(42)
for _ in range(E04_N_EP):
env.reset()
agent.reset()
for t in range(E04_MAX_OBS):
obs = env.observe()
agent.update(obs)
entropy_sums[t] += agent.entropy()
mean_h = entropy_sums / E04_N_EP
results_e04_emp[p] = mean_h
# Theoretical: E[H(b_n)] via numerical sum over observation histories
thy_h = np.array([theoretical_entropy(p, n)
for n in range(1, E04_MAX_OBS + 1)])
results_e04_thy[p] = thy_h
# Note: at n=2 entropy continues to drop even though accuracy = accuracy(n=1)
dh_n2 = thy_h[1] - thy_h[0]
half_life = next((n+1 for n, h in enumerate(mean_h) if h < 0.5), None)
print(f" p={p:.2f} C={channel_capacity(p):.4f}b "
f"H(b₁)={thy_h[0]:.4f} H(b₂)={thy_h[1]:.4f} ΔH(1→2)={dh_n2:+.4f} "
f"steps to H<0.5: {half_life or '>40'}")
fig, axes = plt.subplots(1, 2, figsize=(12, 4.5))
fig.suptitle("E04 — Belief Entropy Decays Monotonically; Accuracy Stagnates at Even n", y=1.01)
ax = axes[0]
for p, color in zip(E04_P_VALS, E04_COLORS):
ax.plot(range(1, E04_MAX_OBS+1), results_e04_emp[p], 'o', color=color, ms=2.5, alpha=0.4)
ax.plot(range(1, E04_MAX_OBS+1), results_e04_thy[p], color=color, lw=2.5,
label=f'p={p:.2f} (C={channel_capacity(p):.3f}b)')
ax.axhline(0.5, color='black', lw=0.8, ls=':', alpha=0.5)
ax.set_xlabel("Observations n")
ax.set_ylabel("Mean belief entropy E[H(b_n)] [bits]")
ax.set_title("Entropy decay: MONOTONE in n\n(line=exact theory, dots=empirical)")
ax.set_ylim(0, 1.05)
ax.legend(fontsize=7)
ax = axes[1]
for p, color in zip(E04_P_VALS, E04_COLORS):
thy = results_e04_thy[p]
valid = thy > 1e-6
ax.semilogy(np.arange(1, E04_MAX_OBS+1)[valid], thy[valid], color=color,
lw=2, label=f'p={p:.2f}')
ax.set_xlabel("Observations n")
ax.set_ylabel("E[H(b_n)] [log scale]")
ax.set_title("Near-exponential entropy decay\n(even-n steps still reduce entropy, unlike accuracy)")
ax.legend(fontsize=7)
plt.tight_layout()
plt.savefig('e04_entropy_decay.png', bbox_inches='tight')
plt.close()
print("Saved: e04_entropy_decay.png")
FINDING 4: Belief entropy and decision accuracy measure different things and behave differently as n grows. Accuracy is pairwise-constant (staircase): P(correct|2k) = P(correct|2k−1). Entropy is MONOTONE: E[H(b_n)] decreases at EVERY step, including even n.
Example at p=0.80: going from n=1 to n=2, accuracy is unchanged (still 0.80), but entropy drops from H(b_1)≈0.7219 to H(b_2)≈0.5395 — a gain of 0.18 bits. The second observation does NOT move the majority vote in expectation, but it does sharpen the posterior distribution. Information accumulates continuously; the ability to translate it into a better binary decision is quantized.
The theoretical curve E[H(b_n)] = Σ_k Binom(k;n,p)·H(σ((2k−n)·logit(p))) matches empirical mean entropy to within simulation noise, confirming the formula. Entropy decays near-exponentially: linear on log scale, with rate ∝ channel capacity C.
E05 — Misspecified channel in a dynamic world¶
Hypothesis: In a STATIC world, any agent with a correct sign on the channel (p_assumed > 0.5) eventually converges to the right answer. But in a DYNAMIC world (p_switch > 0), the agent must TRACK a moving target — over-confidence causes it to over-commit and track state changes too slowly; under-confidence makes it track too cautiously. The correctly specified agent optimally balances commitment and adaptability.
Setup: True p_correct=0.80, p_switch=0.10; p_assumed ∈ {0.60, 0.70, 0.80, 0.90, 0.95}. Agent always knows p_switch correctly. T=200 steps, n_episodes=3000.
print("\nE05: Misspecified channel in a dynamic world")
print("=" * 50)
E05_TRUE_P = 0.80
E05_P_SWITCH = 0.10
E05_ASSUMED_PS = [0.60, 0.70, 0.80, 0.90, 0.95]
E05_T = 200
E05_N_EPISODES = 3000
misspec_colors = ['#9ca3af', '#d97706', '#16a34a', '#0891b2', '#dc2626']
results_e05 = {}
for p_assumed in E05_ASSUMED_PS:
env = DynamicTwoDoorEnv(E05_TRUE_P, E05_P_SWITCH)
agent = DynamicBayesianAgent(p_assumed, E05_P_SWITCH)
np.random.seed(42)
acc = run_dynamic_accuracy(env, agent, E05_T, E05_N_EPISODES)
results_e05[p_assumed] = acc
tag = '✓ correct' if p_assumed == E05_TRUE_P \
else ('↑ over' if p_assumed > E05_TRUE_P else '↓ under')
print(f" p_assumed={p_assumed:.2f} accuracy={acc:.4f} {tag}")
correct_acc = results_e05[E05_TRUE_P]
# Also measure accuracy at different phases of a switch to show tracking speed
# Track: after a state switch, how quickly does each agent adapt?
def track_switch(p_assumed, p_switch, p_true, n_episodes=2000, T_after=30):
"""Accuracy at t=1..T_after steps AFTER a forced state switch."""
env = DynamicTwoDoorEnv(p_true, 0.0) # no random switches
correct_at_t = np.zeros(T_after)
for _ in range(n_episodes):
# Warm up: 50 steps on state=0
env.reset()
env.true_state = 0
agent = DynamicBayesianAgent(p_assumed, p_switch)
agent.reset()
for _ in range(50):
obs = env._observe()
agent.update(obs)
# Now switch state and track recovery
env.true_state = 1
for t in range(T_after):
obs = env._observe()
agent.update(obs)
correct_at_t[t] += int(agent.act() == env.true_state)
return correct_at_t / n_episodes
print("\n Tracking speed after state switch (steps to recover > 0.80 accuracy):")
switch_results = {}
for p_assumed, color in zip(E05_ASSUMED_PS, misspec_colors):
rec = track_switch(p_assumed, E05_P_SWITCH, E05_TRUE_P, n_episodes=2000)
switch_results[p_assumed] = rec
steps_to_80 = next((t+1 for t, a in enumerate(rec) if a >= 0.80), None)
print(f" p_assumed={p_assumed:.2f} steps to 80% recovery: {steps_to_80 or '>30'}")
fig, axes = plt.subplots(1, 2, figsize=(12, 4.5))
fig.suptitle(f"E05 — Misspecified Channel in Dynamic World\n"
f"(true p_correct={E05_TRUE_P}, p_switch={E05_P_SWITCH})", y=1.01)
ax = axes[0]
assumed_vals = E05_ASSUMED_PS
accs = [results_e05[p] for p in assumed_vals]
bars = ax.bar([str(p) for p in assumed_vals], accs,
color=misspec_colors, width=0.6, edgecolor='white')
for bar, acc in zip(bars, accs):
ax.text(bar.get_x() + bar.get_width()/2, bar.get_height() + 0.001,
f'{acc:.4f}', ha='center', va='bottom', fontsize=9)
ax.axhline(correct_acc, color='#16a34a', lw=1.5, ls='--',
label=f'Correct spec ({correct_acc:.4f})')
ax.set_xlabel("p_assumed (agent's belief about sensor quality)")
ax.set_ylabel("Mean accuracy (T=200 steps)")
ax.set_title("Steady-state accuracy under misspecification")
ax.set_ylim(0.75, 0.90)
ax.legend(fontsize=9)
ax = axes[1]
T_after = 30
for p_assumed, color in zip(E05_ASSUMED_PS, misspec_colors):
rec = switch_results[p_assumed]
lw = 2.5 if p_assumed == E05_TRUE_P else 1.5
ls = '-' if p_assumed == E05_TRUE_P else '--'
label = f'p_assumed={p_assumed:.2f}' + (' ← correct' if p_assumed == E05_TRUE_P else '')
ax.plot(range(1, T_after+1), rec, color=color, lw=lw, ls=ls, label=label)
ax.axhline(0.80, color='black', lw=0.8, ls=':', alpha=0.5, label='80% recovery')
ax.set_xlabel("Steps after state switch")
ax.set_ylabel("Accuracy (tracking state=1 after switch from state=0)")
ax.set_title("Tracking speed after a sudden state switch")
ax.legend(fontsize=7)
plt.tight_layout()
plt.savefig('e05_misspecified_channel.png', bbox_inches='tight')
plt.close()
print("\nSaved: e05_misspecified_channel.png")
FINDING 5: In a dynamic world, misspecification of channel quality affects both steady-state accuracy and tracking speed after state changes.
Over-confident agents (p_assumed > p_true) achieve higher certainty per observation, making them faster to commit after a switch — but they also over-commit to stale beliefs. If the state changes again before they've recovered, their over-confidence becomes a liability. At extreme over-confidence (p_assumed=0.95), the agent locks strongly into its current belief and is SLOW to track state changes, even though it reaches certainty quickly when it does update.
Under-confident agents (p_assumed < p_true) update beliefs more cautiously per observation, making them slower to commit but also more responsive to contradicting evidence. At p_assumed=0.60, the agent treats reliable observations as nearly random, staying closer to its prior and being slow to both commit AND track.
Correctly specified agent optimally balances the commitment-adaptability tradeoff. This is the key design insight: for tracking in a dynamic world, the correct model of sensor quality is what achieves optimal steady-state accuracy AND recovery speed.
Summary¶
F.Shannon — Empirical conclusions:
| Experiment | Finding |
|---|---|
| E01 | P(correct|n) = 1−Binom.CDF(⌊(n−1)/2⌋,n,p) confirmed exactly |
| E02 | CLT z-score √n·(2p−1)/√(p(1−p)) collapses curves onto Φ(z/2) |
| E03 | n* ≈ (Φ⁻¹(τ))²·p(1−p)/(p−0.5)²; CLT approximation accurate for p≳0.65 |
| E04 | Entropy decays MONOTONICALLY (even-n steps reduce entropy though not accuracy) |
| E05 | In dynamic worlds, misspecification hurts both steady-state and tracking |
Precision findings: Three quantities (logit(p), C=1−H_b(1−p), and CLT z-rate) are all monotone in p but measure different aspects of channel usefulness. Bayesian updating uses logit(p); Shannon coding theory uses C; normal approximation for n* uses (p−0.5)/√(p(1−p)).
Connection to Area 01 master question: → Information quantity (n · I(O;S)) bounds achievable certainty, but the agent must ALSO model its channel correctly to exploit that information optimally. A wrong model is worse than a conservative one — especially when the world changes.
print("\n" + "=" * 60)
print("Notebook F.Shannon complete.")
print("Outputs: e01_shannon_accuracy_curves.png, e02_cumulative_information.png,")
print(" e03_observations_needed.png, e04_entropy_decay.png,")
print(" e05_misspecified_channel.png")