ONE JOINT GRADIENT. BETTER COORDINATION.

G2MAFTest-Time Gradient Guidance for Multi-Agent Flow Policies

G²MAF uses a frozen centralized critic to coordinate a single refinement of a frozen flow policy’s proposal at test time.

Guowei ZouHaitao WangGuoxin WangZhiquan ChenBeiwen ZhangGuojie WangHejun Wu

Sun Yat-sen University

Explore the methodPaper & code · coming soon
Frozen policies · Test-time coordination · Offline MARL

01 / Overview

Refine the joint action before executing it.

A frozen generative policy can propose a feasible joint action while leaving room for better coordination. G²MAF uses a centralized behavior critic to evaluate the full joint action and provide one joint gradient. A globally normalized update refines the proposal, and projection keeps the executable action within the feasible set.

+9.2%Mean relative gainon MPE at the reported operating points
+8.9%Mean relative gainon SMAC at the reported operating points
20 / 24Settings improvedwith the canonical refinement variant
≈6%Additional latencyin average model inference time
One joint correction recovers part of the value left by the original proposal, within the neighborhood supported by behavior data.
Closing the local coordination gap.One joint correction recovers part of the value left by the original proposal, within the neighborhood supported by behavior data.

02 / Method

Coordinate a single refinement.

Each agent’s gradient depends on the current actions of its teammates. A shared normalization gives the team one update budget. The canonical variant refines the decoded action; the paper also studies guidance within flow trajectories and masked logits for discrete actions.

The critic evaluates the complete joint action. Its gradient is globally normalized, projected, and applied before execution.
One critic, one coordinated update.The critic evaluates the complete joint action. Its gradient is globally normalized, projected, and applied before execution.
01

Propose

Generate a joint action with the frozen multi-agent flow policy.

02

Refine together

Compute one centralized critic gradient and normalize it across all agents.

03

Project & execute

Project the refined proposal onto the feasible action set and execute the joint action.

Inside G²MAF · Before & after

看清 Q 值引导改变了什么。

正在载入真实模型输出…

03 / Videos

Cooperative trajectories in motion.

Play the trajectory videos archived with this paper project.

StarCraft multi-agent micromanagement

SMAC · 3m

Good dataset · episode 31669

SMAC · 2s3z

Good dataset · episode 13583

SMAC · 5m vs 6m

Good dataset · episode 16943

SMAC · 8m

Good dataset · episode 23654

These videos visualize recorded offline dataset trajectories. SMAC movement, allied actions and health follow the recordings, rendered with custom unit artwork; this is not StarCraft client footage or an evaluation of the proposed method. Enemy attack poses are inferred from allied health changes, and a terminal display frame may be appended for recorded wins.

04 / Results

Better joint actions at test time.

The evaluation covers 24 settings across MPE and SMAC. At the reported operating points, the canonical variant improves 20 settings over the same frozen policy without refinement. MPE and SMAC use different score scales, and results are reported separately.

Bold and underlined entries retain the paper’s markings. Scroll wide tables horizontally.

Main results against published baselines

Main results on (a) MPE and (b) SMAC. ``Data'' is the offline data mean, ``Diff'' abbreviates MADiff, and G2MAF denotes the refined frozen backbone at its selected operating point. Bold and underline mark the best and second best method in each row, excluding Data. The G2MAF point estimate is the selected sweep operating point; its ± value is the population standard deviation across K = 1–5 in a separate sweep, not a standard error across seeds. Published baselines retain their source uncertainty conventions.

MPE

TaskQual.DataBCICQTD3+BCCQLOMARDiffMA-SfBCDOM2G2MAF
SpreadExpert96.235.0±2.6104.0±3.4108.3±3.998.2±5.2114.9±2.695.0±5.387.5±7.388.7±6.3115.2±0.9
MedRep9.310.0±3.813.6±5.715.4±5.631.4±7.237.9±6.130.3±2.58.2±4.663.1±9.547.3±2.2
Medium27.031.6±4.829.3±5.539.4±3.634.1±7.247.9±18.964.9±7.751.6±14.278.6±8.159.7±1.7
Random0.1-0.5±3.26.3±3.59.8±4.924.0±9.834.4±5.36.9±3.15.1±3.937.4±11.358.0±9.0
TagExpert89.540.0±9.6113.0±14.4115.2±12.8119.3±14.0123.9±10.5103.0±12.077.4±13.998.2±14.4139.7±1.1
MedRep8.60.9±1.434.5±27.828.7±20.941.7±15.347.1±15.353.9±11.412.7±7.368.2±16.773.5±7.5
Medium43.822.5±1.863.3±20.065.1±29.561.7±23.166.7±23.272.7±9.447.1±17.982.6±18.2112.7±1.1
Random0.01.2±0.52.2±1.36.0±2.111.1±2.811.1±2.84.6±2.611.6±5.129.6±8.146.1±2.9
WorldExpert93.833.0±9.9109.5±22.8110.3±21.3119.8±28.1110.4±25.7109.3±15.497.3±19.199.5±17.1156.4±5.2
MedRep11.22.3±1.512.0±9.117.4±8.119.3±18.342.9±19.519.8±6.29.1±5.965.9±10.645.8±2.0
Medium54.525.3±2.071.9±20.073.4±9.358.6±11.274.6±11.584.7±12.354.2±22.784.5±23.4146.8±11.4
Random0.0-2.4±0.51.0±3.22.8±5.50.6±2.05.9±5.26.1±2.43.1±1.34.1±1.15.0±0.3

SMAC

TaskQual.DataBCICQCQLMA-DTDiffDoFFlow BCMAC-FlowVGM2PG2MAF
3mGood16.516.0±1.018.8±0.619.0±0.319.6±0.719.3±0.519.8±0.220.0±0.019.8±0.219.5±0.720.0±0.1
Medium10.08.2±0.818.1±0.718.9±0.717.2±0.716.4±2.618.6±1.214.7±1.518.0±3.216.9±1.121.4±1.8
Poor4.74.4±0.114.4±1.25.8±0.48.9±0.310.3±6.110.9±1.14.5±0.110.6±2.214.9±1.514.8±0.9
2s3zGood18.318.2±0.419.6±0.319.1±0.819.4±0.115.9±1.218.5±0.819.5±0.119.5±0.519.9±0.121.2±0.1
Medium12.614.3±0.717.2±0.814.3±2.017.4±0.315.6±0.318.1±0.915.1±2.017.6±0.616.5±0.619.1±0.4
Poor6.96.7±0.312.1±0.410.1±0.79.9±0.28.5±1.310.0±1.16.9±0.88.5±0.67.9±0.712.2±0.3
5m_vs_6mGood16.616.6±0.616.3±0.913.8±3.118.0±1.016.5±2.817.7±1.114.7±2.118.6±3.517.6±1.319.4±0.6
Medium12.614.2±0.517.2±0.416.8±3.117.5±0.415.2±2.616.2±0.912.8±0.815.6±1.317.0±0.920.6±0.7
Poor7.57.5±0.29.4±0.410.4±1.08.9±0.38.9±1.310.8±0.37.7±0.89.8±2.110.7±1.112.4±0.6
8mGood16.916.7±0.419.6±0.313.1±6.119.2±0.118.9±1.119.6±0.319.5±0.219.7±0.319.7±0.421.1±0.3
Medium10.110.7±0.518.6±0.516.3±3.118.0±0.516.8±1.618.6±0.818.2±0.819.4±0.618.2±1.619.7±0.4
Poor5.35.3±0.110.8±0.84.6±2.45.1±0.19.8±0.912.0±1.24.9±0.111.5±0.84.9±0.19.5±0.1
Relative gains over the frozen policy without refinement, at the operating points reported in the paper. Uncertainty follows the paper’s figure protocol.
Return gains across 24 settings.Relative gains over the frozen policy without refinement, at the operating points reported in the paper. Uncertainty follows the paper’s figure protocol.

The manuscript has not yet been publicly released.

SMAC success rates

SMAC success rate (↑, our policy only): Frozen backbone vs. G2MAF at each method's selected denoising-step count. Success and shaped episode return measure different outcomes; success is therefore reported separately. Here Δwin=WG2MAF-WFrozen is the percentage-point change in success rate; ΔR is the shaped-return gain.

MapQualityFrozen %G2MAF %ΔwinMapQualityFrozen %G2MAF %Δwin
3mGood95.695.60.05m_vs_6mGood74.478.9+4.4
Medium66.750.0-16.7Medium68.977.8+8.9
Poor22.243.3+21.1Poor13.325.6+12.2
2s3zGood97.895.6-2.28mGood96.798.9+2.2
Medium65.676.7+11.1Medium76.788.9+12.2
Poor0.04.4+4.4Poor0.00.00.0

Inference cost for every setting

Model-side inference latency in milliseconds per decision for every evaluated setting. Base is the frozen backbone; Grad is the direct G2MAF critic-gradient step; G2MAF is the full refined policy.

TaskQualityBase (ms)Grad (ms)G2MAF (ms)×TaskQualityBase (ms)Grad (ms)G2MAF (ms)×
MPESMAC
SpreadExpert142.13.0145.11.023mGood189.93.4179.50.94
MedReplay137.43.0142.71.04Medium175.73.9188.51.07
Medium140.53.2149.11.06Poor221.26.8233.21.05
Random137.32.9143.91.052s3zGood153.84.3181.91.18
TagExpert138.93.1150.71.08Medium164.23.7162.40.99
MedReplay135.63.3147.31.09Poor143.73.4165.41.15
Medium164.55.9168.11.025m_vs_6mGood181.33.6178.70.99
Random141.93.2144.11.02Medium171.33.6163.80.96
WorldExpert136.72.7139.91.02Poor162.93.3176.81.09
MedReplay141.03.6148.01.058mGood165.93.5179.41.08
Medium139.23.2144.61.04Medium176.93.8177.71.00
Random155.55.9156.81.01Poor171.36.3242.91.42
All (24-setting mean)157.93.9167.11.06

Structure, search and local-step controls

Frozen is absolute performance; G2MAF, Best-N, 1-agent and Factor. are paired gains. Best-N uses 16 perturbations; Factor. is MPE-only. These controls use a separate evaluation batch from the main table.

Full-grid mechanism controls and local-step diagnostics over all 24 settings. (a) Paired mean performance gains for the full centralized update and the structure/search controls. (b) Critic-predicted value change and realized update geometry from the same rerun.

Structure and search controls

TaskQualityFrozenG2MAFBest-N1-agentFactor.TaskQualityFrozenG2MAFBest-N1-agentFactor.
MPESMAC
SpreadExpert108.59+0.24+1.19+0.22+0.053mGood19.17+0.01+0.03+0.01--
MedReplay43.29+1.72-2.48-0.03+1.62Medium15.12-1.88-1.70+0.79--
Medium59.35+3.50-4.74+1.24+2.21Poor9.14+1.72+1.04-0.03--
Random36.69+6.84+5.93+1.59+6.302s3zGood19.54+0.32+0.26+0.46--
TagExpert113.02+6.46+7.45+5.05+10.57Medium17.08+1.62+0.75-0.06--
MedReplay60.90+3.77-0.05+0.66+5.99Poor9.87+2.16-0.72+0.57--
Medium94.58+5.24+0.05-0.61+0.055m_vs_6mGood16.97+0.25-0.66-0.05--
Random20.24+10.33+10.00+4.58+12.36Medium15.24+2.32+1.64+2.97--
WorldExpert149.35+2.50+3.59+1.41+2.72Poor11.58-1.22-0.12+0.06--
MedReplay50.11+3.26-0.33-0.33+6.208mGood19.07+0.63+0.40-0.09--
Medium112.72+4.02+2.17-0.54+1.63Medium17.90+1.06+0.40-0.45--
Random0.33+0.98+0.00+0.22+0.65Poor6.62+2.18+1.05+0.33--

Local-step diagnostics

TaskQualityReturnΔQNormBound. frac.TaskQualityReturnΔQNormBound. frac.
MPESMAC
SpreadExpert108.84+3.210.0300.0173mGood19.18+0.030.999--
MedReplay45.02+9.790.0980.030Medium13.24+0.119.076--
Medium62.85+9.850.0980.028Poor10.86+0.2543.610--
Random43.51+5.310.1000.0022s3zGood19.86+0.3210.000--
TagExpert119.53+3.850.0470.081Medium18.70+0.331.941--
MedReplay64.67+54.190.0960.047Poor12.04+0.3127.311--
Medium99.81+16.870.0870.1295m_vs_6mGood17.22+0.081.908--
Random30.52+2.510.0980.042Medium17.56+0.161.923--
WorldExpert151.85+3.890.0820.149Poor10.37+0.171.722--
MedReplay53.37+5.460.0890.0978mGood19.70+0.542.000--
Medium116.74+13.020.0880.148Medium18.95+2.799.835--
Random1.41+0.160.0280.148Poor8.80+1.7630.005--

Held-out step-size transfer and perturbation controls

Full-grid non-oracle step-size transfer and perturbation controls over all 24 settings. (a) Canonical gain, the step selected from the other settings in the same domain, and the independently evaluated held-out gain. (b) Canonical Base and G2MAF gain together with independently measured matched nominal-radius Random and Shuffled effects. MPE values use the corresponding table in the paper OMAR-normalized scale; SMAC values use shaped episodic return. Results across the two panels use different evaluation batches and are compared descriptively.

Held-out step-size transfer

TaskQualityCanonical ΔRLOTO ηLOTO ΔRTaskQualityCanonical ΔRLOTO ηLOTO ΔR
MPESMAC
SpreadExpert+0.300.1-0.763mGood+0.001+0.00
MedReplay+1.400.1+2.12Medium+1.402-1.75
Medium+2.100.1+4.97Poor+0.102-1.61
Random+6.100.1+7.622s3zGood+1.202+1.22
TagExpert+2.500.1+4.76Medium+0.402+0.40
MedReplay+2.600.1+10.38Poor+2.402+1.21
Medium+0.400.1+11.865m_vs_6mGood+1.102+1.08
Random+15.500.1+13.79Medium+1.5010-4.30
WorldExpert-0.200.1+14.72Poor-0.102-0.14
MedReplay+2.200.1+13.148mGood+1.102+1.12
Medium-3.200.1+13.60Medium+2.202+1.95
Random+1.400.1+0.24Poor+2.502+0.67
LOTO positive11/12LOTO positive7/12

Matched-radius perturbation controls

TaskQualityBaseG2MAF ΔRRandom ΔRShuffled ΔRTaskQualityBaseG2MAF ΔRRandom ΔRShuffled ΔR
MPE (OMAR-normalized score)SMAC (episode return)
SpreadExpert114.90+0.30-1.99-1.643mGood20.0+0.0-0.2-1.1
MedReplay45.91+1.40-0.27-0.08Medium20.0+1.4-1.4-2.4
Medium57.60+2.10-8.30-3.42Poor14.7+0.1-8.3-4.2
Random51.91+6.09+2.34-1.102s3zGood20.0+1.2+0.0-0.8
TagExpert137.22+2.50+8.87+5.42Medium18.7+0.4-0.3+1.0
MedReplay70.90+2.59+4.10+5.66Poor9.8+2.4-3.8+1.1
Medium112.31+0.38-5.85-5.755m_vs_6mGood18.3+1.1-0.2+1.1
Random30.61+15.52+3.35+3.68Medium19.1+1.5+0.7-0.6
WorldExpert156.63-0.22+8.59+9.02Poor12.5-0.1+1.0-0.6
MedReplay43.59+2.17+12.17+9.468mGood20.0+1.1+0.3+0.7
Medium150.00-3.15+7.93+12.83Medium17.5+2.2+1.7+1.4
Random3.59+1.41-0.22+0.87Poor7.0+2.5-0.4+1.9

05 / Analysis

Understand the refinement.

The paper examines where to inject the gradient, how the step size affects return, and how critic reliability relates to improvement. The local theoretical result concerns an increase in critic prediction under smoothness and step-size assumptions; realized return is assessed empirically.

The diagnostics compare return gains across dataset qualities and examine the critic’s local ranking reliability.
Dataset quality, headroom, and critic reliability.The diagnostics compare return gains across dataset qualities and examine the critic’s local ranking reliability.
Decoded-action refinement and two trajectory variants are evaluated on the 12 continuous-action MPE settings.
Where should guidance enter the policy?Decoded-action refinement and two trajectory variants are evaluated on the 12 continuous-action MPE settings.

Further experiments

Key ablations and diagnostics.

Where should critic guidance be applied?
Where should critic guidance be applied?. All 12 MPE settings. All and Final are independently evaluated trajectory variants with K = 5; Post is the decoded-action gain at the main-table operating point.
Sensitivity to refinement step size
Sensitivity to refinement step size. Continuous curves use a fixed K = 5 protocol and gains over a matched Base evaluation. MPE uses normalized score; SMAC uses shaped return. Rings mark the best tested discrete step.
Dataset quality, critic reliability and headroom
Dataset quality, critic reliability and headroom. Task-level gains, pairwise critic accuracy and headroom diagnostics. Diagnostic error bars are one standard error across settings; continuous local comparisons use restored MPE states.