How Do Function Vectors Come into Being and Where Do They Live?

Hand-drawn ivory-toned illustration of a pavilion beside a bridge.Shantang Street, Suzhou.

This blog post documents an exploration undertaken from November 2025 to February 2026. The framework explored here has since been expanded to concept learning in an upcoming preprint. The author warmly thanks Alex Warstadt, Mikhail Belkin, Kyle Mahowald, Neil Mallinar, and Sasha Boguraev for their generous time and valuable feedback throughout this exploration.

1. Introduction

Function vectors (Todd et al., 2024) are compact, linear directions in activation space extracted from attention head outputs that causally induce specific task behaviors (such as antonym generation, country-capital retrieval, or translation) when added to the residual stream at inference time.

While the causal efficacy of function vectors in fully trained checkpoints is well established, fundamental questions remain regarding their developmental geometry during pre-training:

  1. How does the function vector evolve across pre-training? When does causal steering emerge, does discrete decision-flipping develop simultaneously with continuous logit push, and is there an abrupt transition during training?
  2. How does the model’s optimization geometry shape this development? How does the function vector relate to the principal directions of loss gradients as the model learns?

The initial intuition behind this investigation was straightforward: during pre-training, do task representations align with the loss landscape as gradients “carve out” functionality; and once the capability is consolidated, does the representation decouple from immediate loss gradients?

Our empirical results reveal a more nuanced picture. By tracing Function Vectors across 19 pre-training checkpoints (from Step 1 to Step 143,000, spanning five orders of magnitude on a logarithmic schedule) in Pythia-410M (evaluated on the antonym task) and Pythia-1B (evaluated across four core in-context learning tasks: antonym, country-capital, english-french, and present-past), and probing optimization geometry using the Average Gradient Outer Product (AGOP) (Radhakrishnan et al., 2024) alongside validation on a converged OLMo-1B model, we map the geometric lifecycle of task representations:

Key Takeaways

2. Setup & Framework

2.1 Models, Tasks, and Extraction

2.2 Gradient Sensitivity and AGOP

To capture the model’s loss landscape geometry, we compute the Average Gradient Outer Product (AGOP) matrix with respect to the attention module output at the final token position under next-token cross-entropy loss as an empirical average over demonstration prompts:

Here, is the scalar cross-entropy loss on the target token. We use the leading orthonormal eigenvectors of (sorted by descending eigenvalue ) as the principal gradient subspace basis for projection and alignment measurements.

We construct two distinct AGOP matrices per layer:

  1. Task AGOP (): Computed over clean task ICL prompts, spanning the task-specific gradient subspace .
  2. Control-data AGOP (): Computed over length-matched uniform random token prompts with identical sequence lengths but arbitrary vocabulary tokens, defining a task-irrelevant baseline subspace .

Both AGOP and FV extraction use only the 128-prompt extraction set; the 200 evaluation prompt pairs remain strictly held-out.

2.3 Tri-Orthogonal Subspace Decomposition

To isolate where causal function resides relative to optimization gradients, we decompose the full Function Vector into three components via sequential projection. The construction strictly guarantees that is orthogonal to the remainder and that is orthogonal to :11 Sequential projection strictly guarantees that and . However, this construction does not generally guarantee that all three components are pairwise mutually orthogonal, because the Task and Control AGOP eigenspaces ( and ) can overlap. On Pythia-1B, the normalized subspace overlap ranges between and , which is small in absolute terms but above the random-subspace expectation of (). On Pythia-410M, we did not track subspace overlap across training.

Figure 1: Tri-orthogonal subspace decomposition framework. The full Function Vector is projected sequentially onto Control-data AGOP () to obtain , then the structural remainder is projected onto Task AGOP () to obtain , leaving the task-orthogonal remainder .The full function vector is decomposed by sequential projection into a control-aligned structural component, a task-parallel component, and a task-orthogonal remainder.
  1. Structural Component (): We first project the raw Function Vector onto the top- eigenspace of Control-data AGOP , yielding . This isolates the portion of the vector aligned with the length-matched random-token control baseline.
  2. Parallel Task Component (): We then take the structural remainder and project it directly onto the top- eigenspace of the original Task AGOP , yielding . Empirically, the task and control eigenspaces are nearly orthogonal in activation space, though this is an empirical property rather than an architectural constraint.
  3. Task-Orthogonal Remainder (): The remaining residual vector . By construction, is strictly orthogonal to the top Task AGOP subspace , and is approximately orthogonal to to the extent that task and control eigenspaces are empirically near-orthogonal in activation space.

To test causal behavior without magnitude confounds, each component is normalized to unit norm and scaled by , where is the layer-average activation norm evaluated at the final token position across evaluation prompts, before injection into the residual stream:

where is the steering intensity (default ).

3. Pre-Training Trajectory

Tracking the Function Vector across 19 pre-training checkpoints reveals that task capabilities do not accumulate linearly. Instead, the model undergoes a sharp, multi-metric transition at Step 1000 (the first ~0.7% of pre-training).

3.1 Causal Metrics and antonym Trajectory on Pythia-410M

When the extracted Function Vector is injected at layer with intensity , we evaluate two complementary causal metrics across 200 held-out evaluation pairs:

Figure 2: Pythia-410M, antonym, γ=1.0. Solid = median across the 24 layers (band = 25th–75th percentile); dashed = the single strongest layer. The two panels use different y-scales.On Pythia-410M antonym, the function vector's logit effect jumps at Step 1000, while discrete top-1 decision flipping matures much later.
Data behind this figure

(target logit)

stepFV medianFV best layerrandom medianrandom best layer
1-0.0344-0.00350.01470.0876
2-0.03310.00420.01480.0876
4-0.03380.01580.01630.0909
8-0.0146-0.00540.02090.0927
16-0.04560.02120.00600.0569
32-0.0676-0.00560.01080.0492
64-0.03230.01000.01250.0671
1280.01430.09860.04160.1270
2560.03810.41310.15770.5342
5120.47270.69040.28810.9092
10002.73964.9398-0.32650.4320
20000.54052.0000-0.44481.2676
40000.71092.5273-0.73630.3809
80003.03225.6045-0.38060.3457
160002.92195.1465-1.2158-0.0410
320003.28716.1953-1.12790.0234
640001.81644.6406-1.5498-0.2246
1200002.66415.1719-1.9824-0.0078
1430001.94144.7617-1.5840-0.0312

(accuracy)

stepFV medianFV best layerrandom medianrandom best layer
10.00000.00000.00000.0000
20.00000.00000.00000.0000
40.00000.00000.00000.0000
80.00000.00000.00000.0000
160.00000.00000.00000.0000
320.00000.00000.00000.0000
640.00000.00000.00000.0000
1280.00000.00000.00000.0000
2560.00000.00000.00000.0000
5120.00000.00000.00000.0000
10000.01000.03000.00000.0100
20000.00000.01000.00000.0100
40000.00500.0400-0.01000.0100
8000-0.01000.0200-0.01000.0000
160000.00500.11000.00000.0100
320000.01500.20000.00000.0100
640000.02500.17000.00000.0100
1200000.07500.41000.00000.0100
1430000.02500.29000.00000.0100

Examining the trajectory in Pythia-410M on antonym reveals several key observations:

  1. The Step 1000 Transition: Read the onset against the random control, not against zero. Up to Step 512 the function vector does no more than an isotropic random vector of the same length (where the random control peak is vs for the function vector), so the minor lift before Step 1000 is not specific to the function vector. Across all 24 layers, the cross-layer median leaps from at Step 512 to at Step 1000. Concurrently, the optimal injection site shifts from the final output layer (Layer 23) to intermediate processing layers (Layer 10). Looking at the single strongest layer per checkpoint, the peak effect leaps by (from at Layer 23 at Step 512 to at Layer 10 at Step 1000). Note that is the peak at Step 1000; the global trajectory peak across all checkpoints reaches at Layer 9 at Step 32,000.
  2. Delayed Decision-Flipping Maturation: While directional logit force emerges at Step 1000, discrete next-token classification flipping develops much later. At Step 1000, at Layer 10 is only (2 out of 200 evaluation prompt pairs). From Steps 1000 to 8000, the gap over random control remains small and non-monotonic (0.020 to 0.030, returning to 0.000 at Step 2000). The first checkpoint where the gap reaches and remains sustained above that level is Step 16,000 ( peak at Layer 9 vs for random control), eventually reaching a peak of at Step 120,000.

3.2 Multi-Task Replication on Pythia-1B

Does this early transition generalize beyond the antonym task? Evaluating Pythia-1B across 19 checkpoints on four distinct tasks demonstrates that this two-stage causal maturation is a general property of pre-training dynamics.

Figure 3: Pythia-1B, four tasks, γ=1.0, k=20. Each line is the strongest layer at that checkpoint; solid = function vector, thin dashed = a random vector of the same length in the same task.Across four Pythia-1B tasks, logit effects separate from random control at Step 1000, while top-1 separation emerges later and depends on the task.
Data behind this figure

(target logit)

stepantonym FVantonym randcountry-capital FVcountry-capital randenglish-french FVenglish-french randpresent-past FVpresent-past rand
10.01360.05980.04410.02890.01350.02090.04030.0641
20.05030.05970.02960.02890.01360.02080.02700.0642
40.05480.05940.06520.02880.01380.02070.03010.0638
80.03800.04260.03540.02800.02120.01480.04960.0493
160.02230.01790.03400.03960.01720.02930.05060.0338
320.01570.02600.02300.04980.00390.0221-0.00020.0389
64-0.00180.0256-0.01380.04930.02370.02400.03920.0598
1280.21000.18970.00070.10270.05630.06120.07350.1693
2560.61620.57030.50970.33730.48780.38820.54200.5776
5120.28030.44820.53520.16410.30470.3252-0.03470.4773
10002.02930.00593.1914-0.24612.71780.21882.46680.1660
20004.95410.68362.54690.21881.39990.66505.55960.7080
40002.2441-0.64453.1562-0.79300.8643-0.29106.0605-0.5176
80003.0781-0.72072.1484-1.07810.5557-0.18957.7617-0.2324
160004.1816-0.55474.1484-1.40621.9619-0.14459.56840.0098
320006.2930-0.67583.7773-0.82032.9512-0.02648.1855-0.6074
640004.50980.65433.7695-0.50394.73240.37507.5078-0.2617
1200004.15430.30273.60940.78124.65820.21888.7031-0.3398
1430004.51950.47663.69140.90234.26760.19738.0625-0.5078

(accuracy)

stepantonym FVantonym randcountry-capital FVcountry-capital randenglish-french FVenglish-french randpresent-past FVpresent-past rand
10.00000.00500.00000.00000.00000.00000.00000.0000
20.00000.00500.00000.00000.00000.00000.00000.0000
40.00000.00500.00000.00000.00000.00000.00000.0000
80.00000.00000.00000.00000.00000.00000.00000.0000
160.00000.00000.00000.00000.00000.00000.00000.0000
320.00000.00000.00000.00000.00000.00000.00000.0000
640.00000.00000.00000.00000.00000.00000.00000.0000
1280.00000.00000.00000.00000.00000.00000.00000.0000
2560.00000.00000.00000.00000.00000.00000.00000.0000
5120.00000.00000.00000.02000.00000.00000.00000.0000
10000.01000.00500.00500.01000.01000.01500.01500.0100
20000.05500.02000.02000.04500.00500.02000.08000.0100
40000.00500.00500.10500.01000.01000.00500.01500.0100
80000.07500.00000.09500.01000.00000.00000.13000.0050
160000.21500.00500.08000.00500.02000.01000.10500.0000
320000.37500.01000.10500.02000.04500.01500.07500.0150
640000.28000.04000.11000.01500.09000.01500.07500.0400
1200000.10500.02500.1400-0.00500.09500.02000.06500.0000
1430000.18500.02500.20500.03500.07000.02000.06000.0500

4. Optimization Geometry

How does the Function Vector align with the model’s loss landscape during training? We project the extracted FV onto the top- eigenspaces of Task AGOP () and Control-data AGOP ().

4.1 Pythia-410M Alignment Trajectory

On Pythia-410M (), we track the alignment of the Function Vector with AGOP eigenspaces across training steps compared against both an empirical isotropic random-vector null control and the analytic chance baseline (representing the root-mean-square alignment for an isotropic random vector; see Appendix for proof sketch).

Figure 4: Pythia-410M, antonym, k=128. Line = median across the 24 layers; band = 25th–75th percentile. Solid = the function vector, dashed = an isotropic random vector of the same length. Blue and orange measure alignment with the task AGOP subspace, green and amber with the control AGOP subspace. The chance level is , the root-mean-square alignment a randomly drawn direction achieves with any 128-dimensional subspace of a 1024-dimensional space.On Pythia-410M, function vector alignment with the task AGOP subspace departs from baseline only around Steps 1000–2000 and returns to chance by convergence.
Data behind this figure
stepFV vs taskrandom vs taskFV vs controlrandom vs controlFV - randomratio FV/randomFV: control - taskrandom: control - task
10.38000.35010.38780.35950.02991.08540.00780.0094
20.35650.34970.37090.36410.00681.01960.01440.0144
40.34470.34530.34780.3442-0.00060.99820.0031-0.0012
80.34910.35000.33970.3430-0.00090.9973-0.0094-0.0070
160.36180.35210.33830.36280.00971.0275-0.02350.0107
320.37620.34790.38490.35920.02831.08150.00870.0113
640.33450.35730.35940.3541-0.02280.93630.0249-0.0032
1280.36220.35310.35250.35260.00911.0258-0.0096-0.0005
2560.35320.35280.35450.34640.00031.00090.0013-0.0065
5120.37910.35530.38610.35290.02381.06690.0070-0.0024
10000.47880.36630.51530.38480.11251.30700.03650.0185
20000.45470.36320.48870.37690.09151.25190.03400.0137
40000.37030.35670.41730.36200.01361.03810.04690.0053
80000.36730.35490.41610.36550.01241.03490.04880.0107
160000.36690.35370.41860.36030.01321.03740.05170.0066
320000.35260.35440.40900.3646-0.00170.99510.05640.0102
640000.36120.35570.35840.35160.00551.0155-0.0028-0.0041
1200000.35610.35460.36290.35450.00151.00430.0068-0.0001
1430000.36380.35910.35750.35300.00471.0131-0.0063-0.0061

The measured random control confirms the chance baseline empirically, sitting 2–4% above the analytic line. In the function vector trajectory:

4.2 Multi-Task Alignment on Pythia-1B

On Pythia-1B (), we evaluate AGOP alignment across all four tasks against the analytic chance baseline .

Figure 5: Pythia-1B, four tasks, k=20. Line = median across the 16 layers; band = 25th–75th percentile. The chance level 0.099 is √(k/d)=√(20/2048).On Pythia-1B, AGOP alignment across four tasks peaks at Steps 1000–2000, dips mid-training, and rebounds to settle above chance.
Data behind this figure
stepantonymantonym xchancecountry-capitalcountry-capital xchanceenglish-frenchenglish-french xchancepresent-pastpresent-past xchance
10.10451.05750.09090.91980.09320.94320.10731.0853
20.09120.92320.08780.88830.09310.94240.11431.1570
40.10381.05080.08420.85180.09320.94280.10141.0257
80.09200.93120.09140.92490.09130.92390.08950.9061
160.09450.95650.10841.09660.09080.91920.08620.8726
320.11121.12490.08410.85080.12831.29810.10401.0519
640.09250.93620.07830.79280.10311.04350.08240.8342
1280.10151.02690.13271.34260.07340.74320.09891.0007
2560.11041.11760.12091.22300.09580.96940.10611.0737
5120.13901.40660.11611.17510.14591.47630.13761.3923
10000.19441.96730.16511.67080.17001.72060.20712.0959
20000.13281.34360.11261.13930.22402.26680.16661.6859
40000.11451.15890.15221.54000.20622.08690.15091.5271
80000.11121.12550.12401.25520.12541.26920.14811.4984
160000.13911.40750.10551.06740.13351.35120.14821.4994
320000.12211.23580.11881.20180.17081.72820.15921.6113
640000.14681.48520.14231.43970.15571.57590.14581.4755
1200000.16411.66030.13171.33280.16161.63550.16401.6600
1430000.16061.62490.14401.45730.19171.93980.15611.5792

All four tasks scatter around chance until Step 512; early excursions (such as country-capital at Step 128 reaching chance) are transient and not sustained. From Step 512, all four tasks rise to a peak at Steps 1000–2000 at chance (english-french peaking at Step 2000 at chance; antonym, country-capital, and present-past peaking at Step 1000 at chance). The trajectories then experience dips that vary by task (troughing at Step 8000 for antonym and english-french, Step 16,000 for country-capital, and a shallow dip at Step 64,000 for present-past), before recovering in late training to settle consistently above chance baseline at convergence ( chance, ranging from to ).

The two models used different subspace truncation sizes ( for Pythia-410M, for Pythia-1B), so only the qualitative shapes and each model’s own endpoint relative to its baseline are comparable, not the raw peak heights. Unlike the 410M evaluations, these 1B runs did not record an empirical random-vector control, so the baseline shown is the analytic chance line.

4.3 Divergent Alignment Dynamics

Across the two model scales, the alignment between Function Vectors and task gradient eigenspaces shows distinct developmental trajectories:

5. Subspace Energy

To understand how the Function Vector distributes across the decomposed subspaces, we examine energy allocation across pre-training. For each component , its energy fraction is defined as its normalized squared norm, , such that the three components strictly sum to . For a single projection, this equals the squared cosine alignment with the subspace. Fixing on Pythia-410M (), the chance expected energy fraction for an isotropic random vector projected onto the -dimensional control subspace is exactly . For the task-parallel component , the expected energy fraction is , where measures subspace overlap. For the task-orthogonal remainder (), the expected energy fraction is , achieving the lower bound of when the two subspaces are strictly orthogonal ().

Figure 6: Energy decomposition across three subspaces during pre-training for Pythia-410M (evaluated on the antonym task). Shaded bands represent the interquartile range across network layers within a single checkpoint. Dashed reference lines denote random-vector baselines: represents the exact expectation for and an upper reference for , while marks the theoretical lower bound for .Away from Step 1000 the three components sit near random baselines; at Step 1000 the control-aligned structural component surges while the task-orthogonal remainder drops.
Data behind this figure
stepV_structV_parallelV_perp
10.15040.11250.7345
20.13760.09750.7644
40.12100.09260.7893
80.11540.09710.7890
160.11450.11140.7763
320.14810.11660.7353
640.12910.09470.7769
1280.12430.10860.7718
2560.12560.10040.7731
5120.14910.10640.7457
10000.26560.11900.6195
20000.23890.11610.6385
40000.17410.10670.7173
80000.17310.10370.7139
160000.17540.10620.7266
320000.16730.09820.7361
640000.12850.10670.7639
1200000.13170.10540.7579
1430000.12780.10810.7651

6. Low-Rank Concentration

How deeply does the Function Vector align with the Task AGOP eigenspace? By sweeping the truncation rank on Pythia-410M () at the final checkpoint (Step 143,000), we measure the Enrichment Ratio, defined as the energy fraction of in the top- Task AGOP subspace divided by the random chance baseline :

Figure 7: Pythia-410M, antonym, final checkpoint (step 143000). Band = 25th–75th percentile across layers.Function vector alignment is concentrated in the leading 1 to 4 task AGOP eigenvectors and decays to chance by k=128.
Data behind this figure
kmedianp25p75layer 10
13.44850.79055.943912.6146
22.54871.11454.03356.3251
42.35431.57713.50853.5061
81.81701.43092.52852.1695
161.31421.12662.04462.0254
321.25751.02811.41091.3866
640.99250.89921.12791.2126
1280.86510.82600.93700.9170

The Function Vector is concentrated along the leading 1 to 4 eigenvectors of the Task AGOP matrix:

7. Component Causal Efficacy

When we isolate , , and , scale each to unit layer norm, and inject them individually at inference, how do their causal effects compare to the intact Function Vector?

Figure 8: Pythia-410M, antonym, k=128, γ=1.0. Solid with markers = the strongest layer at that checkpoint; faint dotted in the same colour = the median layer for the same component.The isolated task-parallel component beats the intact function vector in late-training top-1 flipping, while the intact vector keeps the largest logit push.
Data behind this figure

(target logit)

stepfull FVV_parallelV_perpV_struct
1-0.00350.0779-0.02240.0756
20.00420.05110.00350.0991
40.01580.0632-0.00700.1173
8-0.00540.0506-0.00780.0969
160.02120.05890.00040.0160
32-0.00560.0432-0.00690.0415
640.0100-0.00230.01770.0628
1280.09860.12980.09940.0784
2560.41310.52640.39650.4316
5120.69041.05960.50980.9297
10004.93982.26990.68844.7914
20002.0000-0.21881.40822.9883
40002.52730.70311.30471.6992
80005.60453.61623.45215.2295
160005.14654.58013.45904.4785
320006.19535.78913.91415.9141
640004.64063.59382.19533.8438
1200005.17193.86723.20313.4062
1430004.76173.02733.15233.1914

(accuracy)

stepfull FVV_parallelV_perpV_struct
10.00000.00000.00000.0000
20.00000.00000.00000.0000
40.00000.00000.00000.0000
80.00000.00000.00000.0000
160.00000.00000.00000.0000
320.00000.00000.00000.0000
640.00000.00000.00000.0000
1280.00000.00000.00000.0000
2560.00000.00000.00000.0000
5120.00000.00000.00000.0000
10000.03000.01000.00000.0400
20000.01000.01000.00000.0200
40000.04000.02000.01000.0200
80000.02000.00000.00000.0400
160000.11000.20000.01000.0800
320000.20000.32000.05000.1400
640000.17000.25000.05000.0200
1200000.41000.37000.19000.1200
1430000.29000.31000.20000.0500

Every component is injected at the same length (rescaled to times the layer’s mean activation norm), testing directional efficacy rather than magnitude:

7.1 Metric Divergence on Pythia-410M and OLMo-1B

Decomposing the Function Vector reveals a divergence between continuous logit steering and discrete classification success:

7.2 Cross-Task Component Advantage on Pythia-1B

To test whether isolating improves discrete accuracy across multiple tasks, we evaluate the difference in between the isolated task-parallel component and the intact Function Vector across 19 checkpoints in Pythia-1B ().

Figure 9: Pythia-1B, four tasks, γ=1.0, k=20. Each line is the difference between two Δtop-1 accuracies measured at each checkpoint’s strongest layer: the isolated task-parallel component minus the intact function vector. Above the zero line the isolated component flips more decisions; below it the intact vector does. Both are injected at the same length.The isolated component's top-1 advantage over the intact vector is task-dependent on Pythia-1B: stable for country-capital and present-past, late-emerging for antonym, and absent for english-french.
Data behind this figure
stepantonym V_parantonym full FVantonym differencecountry-capital V_parcountry-capital full FVcountry-capital differenceenglish-french V_parenglish-french full FVenglish-french differencepresent-past V_parpresent-past full FVpresent-past difference
10.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.0000
20.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.0000
40.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.0000
80.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.0000
160.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.0000
320.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.0000
640.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.0000
1280.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.0000
2560.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.0000
5120.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.00000.0000
1000-0.01000.0100-0.02000.01500.00500.0100-0.01500.0100-0.0250-0.02500.0150-0.0400
20000.01500.0550-0.04000.07000.02000.05000.00000.0050-0.00500.08000.08000.0000
40000.02500.00500.02000.06500.1050-0.04000.01500.01000.00500.11500.01500.1000
80000.10000.07500.02500.10000.09500.0050-0.00500.0000-0.00500.17000.13000.0400
160000.16000.2150-0.05500.20500.08000.12500.02500.02000.00500.19000.10500.0850
320000.36500.3750-0.01000.29500.10500.19000.04000.0450-0.00500.16000.07500.0850
640000.32500.28000.04500.10500.1100-0.00500.08000.0900-0.01000.14000.07500.0650
1200000.12000.10500.01500.15500.14000.01500.06500.0950-0.03000.12000.06500.0550
1430000.25500.18500.07000.29000.20500.08500.04500.0700-0.02500.07500.06000.0150

The component advantage in Pythia-1B shows substantial task-dependent heterogeneity:

Furthermore, on continuous logit steering (), Pythia-1B exhibits dynamics distinct from Pythia-410M: on Pythia-1B, frequently exceeds the intact Function Vector in raw logit lift (e.g., at Step 64,000 across all four tasks). The claim that the intact Function Vector consistently produces larger logit shifts is specific to Pythia-410M and OLMo-1B (both evaluated on antonym).

7.3 Margin Selectivity and Classification Mechanics

Why does flip more discrete decisions despite producing a smaller raw target logit increase on Pythia-410M? An inspection of prediction margins () reveals how the three components navigate the trade-off between push magnitude and selectivity:

This pattern is corroborated in probability space by target-versus-rest log-odds (): at Step 32,000 (Layer 9), achieves compared to for and for .

These measurements show that discrete classification flipping requires both sufficient target push and selectivity against competitor elevation: pushes strongly but non-selectively, is selective but weak in raw margin at Step 32,000, and only satisfies both criteria at both checkpoints examined (note: tracks the maximum logit among non-target tokens, and the identity of this maximizing token may change before and after intervention).

8. Limitations

We explicitly state the empirical boundaries of this study:

  1. Model and Checkpoint Scope: Full developmental trajectories across 19 pre-training checkpoints are analyzed on Pythia-410M and Pythia-1B. Validation on OLMo-1B is restricted to its single converged checkpoint (step1454000-tokens3048B), as intermediate pre-training checkpoints were not available. On Pythia-410M, the control-aligned component plays a non-trivial causal role in the middle training stage that it does not retain at convergence; because OLMo-1B provides only a converged checkpoint, this mid-training trajectory has not yet been tested there.
  2. Layer Variance vs. Statistical Error: Shaded bands in our figures represent the interquartile range (25th–75th percentile) across network layers within a single checkpoint run, illustrating how intervention efficacy varies across injection layers. They do not represent sampling variance across evaluation prompts or cross-seed training variance. The study plots no statistical error bars.
  3. Subspace Hyperparameters: The multi-task component advantage evaluation on Pythia-1B was conducted with a single random seed and fixed subspace rank .

9. Conclusion

By observing Function Vector causal efficacy throughout pre-training and decomposing vectors against empirical loss gradient eigenspaces (AGOP), we map the developmental geometry of linear task representations:

Appendix: Expected Energy Fraction Proof Sketches

The theoretical baselines for subspace energy and alignment derive from projection properties of isotropic random vectors:

  1. Energy Fraction in a -Dimensional Subspace (): For a standard Gaussian random vector , the total energy . Projecting onto an arbitrary -dimensional subspace yields projected energy , while the orthogonal component . Since and are independent, the energy fraction follows a Beta distribution . Its expected value is:

    The root-mean-square cosine alignment is .

  2. Task-Parallel Component ():

    where is the subspace overlap. Thus serves as an upper bound, with equality achieved if and only if the subspaces are strictly orthogonal ().

  3. Task-Orthogonal Remainder ():

    Thus serves as an exact lower bound, achieved if and only if the task and control subspaces are strictly orthogonal ().

References

Cite this post

@misc{wang2026functionvectors,
  author       = {Yupei Wang},
  title        = {How Do Function Vectors Come into Being and Where Do They Live?},
  year         = {2026},
  month        = {aug},
  doi          = {10.5281/zenodo.22109632},
  howpublished = {\url{https://ypwang.one/Blog/fv-agop-dev-interp/}},
  urldate      = {2026-08-25},
  note         = {Blog post},
}