# The Coach Selection Index — India **CSI-IN v1.0 · A buyer-operated, platform-independent method for ranking executive coaches** Specification date: 16 September 2026 · Licence: CC BY 4.0 · Spec SHA-256: `54e5822d…f54a5` Written and funded by Nirvedha Executive Coaching Solutions · hosted independently · author ranked, but never codes himself (see §12) --- ## 0. The one-paragraph version India has roughly no honest way for a buyer to compare executive coaches. What passes for ranking is either an award, a directory that sells position, or a media list built on visibility. CSI-IN inverts the direction of the instrument: the **buyer** operates it, not the seller. A buyer states two things — who is paying, and at what leadership level — or drags seven sliders if the brief is unusual, and receives a weighted score out of 100 for each coach under consideration, computed from 21 countable indicators across 7 pillars, every point of which must be backed by evidence the buyer can open. The weights are not our opinion. They are derived from what 38 corporate sponsors and 54 coachees in South and Southeast Asia said they actually select on, published cell by cell with their provenance. The arithmetic runs outside the language model, so the same inputs produce the same number on Claude, ChatGPT, Gemini, a spreadsheet, or a sheet of paper. --- ## 1. Why the existing rankings fail the buyer There is no shortage of "top coach" lists. There is a shortage of lists a buyer can use. **They rank firms, not coaches.** Vistage's 2026 list evaluated 70+ executive coaching *firms* on five weighted factors — Coach Experience & Quality 30%, Program Structure & Methodology 25%, Peer Learning Integration 20%, Proven Business Results 15%, Global Reach & Accessibility 10%. A CHRO in Hyderabad hiring one coach for one CXO cannot act on that. Worse, "Peer Learning Integration" at 20% is a criterion on which the publisher's own peer-advisory product scores maximally. When the ranker sells in the category it ranks, the weight vector is the product pitch. **They rank visibility.** Thinkers50 — the most respected instrument in the adjacent category — is explicit that it measures *viability* and *visibility* across ten criteria. That is the right design for ranking ideas. It is precisely the wrong design for hiring a coach, and we have the data to say so: in the ICF South and Southeast Asia study, **visibility ranked eighth of nine factors among corporate sponsors, at 2.6%.** A ranking weighted on reach measures marketing budget. CSI-IN excludes visibility as a parameter entirely (§5.4). **They are unauditable.** "Proprietary assessment methodology" is the standard phrase. It means the buyer cannot check the arithmetic, cannot see which coach lost on which criterion, and cannot tell whether the list is a ranking or a rate card. **They are static where the buyer is specific.** A single national list assumes the best coach for a first-time manager in Pune is the best coach for a promoter-chairman in Chennai. The entire premise is wrong. There is no best coach in India. There is only the best-evidenced coach for a stated buyer at a stated level — which is why CSI-IN produces **twelve** different weight vectors, not one. --- ## 2. Design principles These six commitments constrain every later decision. Where a principle and a convenience collided, the principle won. 1. **The buyer holds the instrument.** Coaches do not submit to CSI-IN. Buyers run it on coaches. No coach can pay to be listed, pay to rank, or pay to be removed. 2. **Evidence beats assertion, always.** Every point awarded is tied to an evidence tier. An unverifiable claim scores 70% of a verified one; an absent claim scores zero however loudly it is asserted (§6). 3. **No relative normalisation, ever.** A coach's score is computed against fixed external anchors, never against the other candidates. This single decision is what makes the index reproducible (§9.1). 4. **The model does not judge; it counts.** Every indicator is a countable or documentary test. Not "how clear is their approach, 1–10" but "is there a written process of a page or more: yes or no." Adjectives are where reproducibility goes to die. 5. **Non-compensatory floors.** A weighted sum alone lets a coach buy the top of the table on cheap dimensions. Five gates screen before any scoring happens (§7). 6. **Publish the uncertainty.** A score built on thin evidence is reported with a wide confidence interval, and coaches whose intervals overlap are not separated into false ranks (§8). --- ## 3. The COACHES framework — seven pillars | | Pillar | What it asks | |---|---|---| | **C** | Context Fit | Has this coach worked in my sector, at my level, and have they themselves led? | | **O** | Outcome Evidence | Can they show change that somebody other than themselves confirmed? | | **A** | Authenticity & Chemistry | Will they sit in a real chemistry session, and do their claims survive checking? | | **C** | Credential & Ethics Floor | Credential, supervision, CPD, ethics, indemnity. | | **H** | Hours & Practice Depth | Logged hours, years, and whether clients come back. | | **E** | Explicit Method | Is the approach written down and testable before I buy? | | **S** | Straight Pricing | Is the fee knowable, justified, and are the commercial terms written? | Each pillar carries **three indicators**, each scored **0–4** against published anchors — 21 indicators in total. The full anchor set is in `csi-spec-v1.0.json`. Two examples: > **HRS1 · Documented coaching hours** — 0: under 100 · 1: 100–499 · 2: 500–1,499 · > 3: 1,500–2,999 · 4: 3,000 or more > **OUT1 · Measurement design** — the strongest outcome measure the coach agrees *in writing > before the engagement starts*. 0: no written measure · 1: coachee self-report only · > 2: + structured stakeholder feedback · 3: + pre/post 360 or psychometric · > 4: + business-impact analysis Indicators within a pillar are **equally weighted**. This is deliberate. Dawes's work on improper linear models showed that unit weights routinely match or beat statistically optimised weights on fresh data, because optimised weights overfit the sample they were tuned on. With 38 sponsors in the underlying study, fitting sub-weights would be false precision. Equal weighting is both more robust and easier to audit. ### 3.1 The one subjective input, and whose it is **AUTH2 — buyer-rated chemistry** is the only judgement in the model, and it belongs to the buyer, after they have sat in the session. No AI ever scores it. No coach ever scores it. This matters: chemistry is the second-most-cited coachee selection factor at 61.1%, so leaving it out would be dishonest — but letting a language model infer it from a website would be worse. ### 3.2 What was moved out of the coach's score The ICF study's strongest success factors include **coachee readiness** (67.6% among sponsors, 57.4% among coachees), **organisational culture** (35.1%) and **manager support** (21.6%). These are real, and they are not attributes of the coach. Scoring a coach on them would reward coaches for the luck of their clients. They are therefore relocated, not discarded, into the **Engagement Readiness Check** — a separate five-question score the buyer runs on *themselves*. It never enters any coach's CSI. A buyer who scores badly on readiness is told plainly that changing coaches will not fix it. --- ## 4. The buyer template — two questions, twelve vectors The buyer answers two questions. Nothing else changes the weights. **Q1 · Who is paying?** Personal (self-funded) · Corporate (organisation-sponsored) **Q2 · What leadership level?** | | Level | |---|---| | L1 | Emerging leader / first-time manager | | L2 | Mid-level manager | | L3 | Senior manager / function head | | L4 | SBU head / Director / VP | | L5 | CXO / C-suite | | L6 | Board / Chair / Promoter / Founder-CEO | 2 × 6 = **12 published weight vectors**. All twelve are in §5.3 and in the spec file. A buyer can read the exact numbers that will be applied to their shortlist *before* they run it. ### 4.1 Tuning the weights to an unusual brief Twelve vectors cover the ordinary cases. They do not cover a turnaround, a first-generation promoter handing over to professional management, a leader relocating across cultures, or a regulated-industry appointment where ethics carries more than the survey average. So the calculator exposes the seven pillar weights as **sliders**. The sliders open on the published vector for the buyer's profile — exactly, to one decimal place — and every move is renormalised so the weights still total 100. The ranking rebuilds live. Three rules keep this from becoming a way to manufacture a preferred answer: 1. **The floor still applies.** No pillar can be pushed below 2%. A buyer cannot switch Outcome Evidence off and rank on price alone. 2. **Deviation is displayed, not hidden.** The interface shows each pillar's movement against the published baseline, and the total displacement in percentage points. 3. **Tuned runs are labelled and exported.** The result carries `weight_mode: "buyer-tuned"` alongside both the tuned vector and the published one it departed from. Reproducibility is unaffected, and this is worth being precise about. The model is deterministic in *(codes, tiers, weights)*. Weights are an **input**, not a judgement made inside the model — so a tuned run is exactly as reproducible as a published one, provided the weight vector travels with the result. It does. What a tuned run loses is not determinism but *authority*: the published vector can claim to represent what 38 sponsors and 54 coachees said; a tuned vector represents what one buyer wanted. Both are legitimate. Only one of them should be published as a market ranking. --- ## 5. How the weights were derived This is the part most rankings do not show. Ours is shown in full, and the code that produces it is `csi_reference.py`. ### 5.1 The evidence base Weights are derived from **ICF *Coaching Forward: State of Executive Coaching in South and Southeast Asia, Trends 2025–26*** (Kirtane, Kaur, Saksena, Bhowmick et al.), n = 397 — 306 coaches, 54 coachees, 37–38 organisational sponsors, across eight countries, India-majority in every segment (coaches ≈66%, coachees ≈80%, sponsors ≈92%). Fieldwork 15 August – 11 November 2025. Ethical approval: Henley Business School, University of Reading. Analysis in SPSS v27, non-parametric methods, 5% threshold, inter-rater agreement above 85%. It is, to our knowledge, the best buyer-side dataset that exists for this market. It is also small on the sponsor side (n=38), and §13 says what that costs us. Two exhibits carry the weights: **Exhibit 5 — sponsor selection factors (n = 38, multiple response)** | Factor | % | |---|---| | Experience level | 68.4 | | Industry knowledge | 52.6 | | Clarity on approach | 47.4 | | Authenticity | 42.1 | | Accreditation or certification | 34.2 | | References | 21.1 | | Fees | 21.1 | | Visibility | 2.6 | | Executive presence and articulation | 2.6 | **Exhibit 17 — coachee selection factors (n = 54, multiple response)** | Factor | % | |---|---| | Coach expertise | 87.0 | | Chemistry / sample sessions | 61.1 | | Authenticity | 59.3 | | Coachee readiness | 57.4 | | Experience | 55.6 | | Credentials | 22.2 | | Fees | 20.4 | ### 5.2 The five derivation steps **Step 1 — Map items to pillars, and sum.** Respondents cast up to three picks each. A pillar's raw weight is its **share of all picks cast**, so where two survey items map to one pillar their values are added. Example: corporate AUTH = authenticity 42.1 + executive presence 2.6 = 44.7. One item required splitting. Exhibit 17's "coach expertise" (87.0) spans two constructs that Exhibit 5 keeps separate — depth of craft and clarity of method. It is apportioned **50/50** to Hours (43.5) and Explicit Method (43.5). This is a declared judgement, flagged `SPLIT` in every table. **Step 2 — Remove what is not a coach attribute.** Coachee readiness, organisational culture and manager support are excluded (§3.2). Visibility is excluded on principle (§5.4). **Step 3 — Impute what the survey did not ask.** Exhibit 17 never offered coachees "industry knowledge" or "references" as options. Their absence is an artefact of the instrument, not a finding that personal buyers do not care. Missing pillars take the **other segment's normalised share**, and the vector is renormalised. Every imputed cell is flagged `IMP`. Two of fourteen base cells are imputed; both are on the personal side. **Step 4 — Apply the declared evidence adjustment, Δ.** Exhibit 5 offered sponsors "references" as the only proxy for demonstrated outcome, which understates how much the same sponsors say evidence matters: elsewhere in the same study, sponsors weight business-impact analysis at **41%** and pre/post 360 at **41%**, against coaches' own practice of 19% and 17%. The report's own conclusion is that measurement is "the price of entry." Δ moves **6 percentage points from Hours to Outcome Evidence**. It is one number, it is published, it is tunable, and §10 shows what happens across Δ = 0…12. Set Δ = 0 to run CSI-IN on pure revealed preference. **Step 5 — Apply the leadership-level tilt.** This layer is **declared judgement, not survey- derived**, and is labelled as such everywhere it appears. The ICF study does not segment selection criteria by leadership level, so nothing empirical is available. Rather than pretend otherwise, the tilt is constrained so it cannot do much damage: **every row sums to zero, and no single entry exceeds ±5 percentage points.** After tilting, no pillar may fall below a 2.0% floor; any deficit is reclaimed proportionally from pillars with headroom. | Level | CTX | OUT | AUTH | CRED | HRS | MTH | PRC | Rationale | |---|---|---|---|---|---|---|---|---| | L1 | −1 | −2 | +3 | 0 | −4 | +2 | +2 | Fit and teachability dominate; CXO lineage is not yet relevant; budgets tightest. | | L2 | −1 | −1 | +2 | 0 | −2 | +1 | +1 | As L1, moderated. | | L3 | +2 | +1 | 0 | −2 | 0 | +1 | −2 | Functional depth begins to bite; credential stops differentiating. | | L4 | +3 | +3 | −1 | −3 | +1 | 0 | −3 | Context and demonstrated outcomes become the decision. | | L5 | +2 | +4 | 0 | −5 | +4 | −1 | −4 | Depth with senior populations and hard evidence peak; every credible candidate already holds a credential. | | L6 | +3 | +2 | +1 | −5 | +5 | −1 | −5 | Governance context and depth peak; confidentiality is handled by the gate, not the score; fee is immaterial. | ### 5.3 The twelve published weight vectors **Personal (self-funded)** | Level | CTX | OUT | AUTH | CRED | HRS | MTH | PRC | |---|---|---|---|---|---|---|---| | L1 | 13.5 | 9.8 | 34.4 | 5.8 | 15.9 | 13.3 | 7.3 | | L2 | 13.5 | 10.8 | 33.4 | 5.8 | 17.9 | 12.3 | 6.3 | | L3 | 16.5 | 12.8 | 31.4 | 3.8 | 19.9 | 12.3 | 3.3 | | L4 | 17.5 | 14.8 | 30.4 | 2.8 | 20.9 | 11.3 | 2.3 | | L5 | 16.2 | 15.5 | 30.8 | 2.0 | 23.4 | 10.1 | 2.0 | | L6 | 17.0 | 13.4 | 31.4 | 2.0 | 24.1 | 10.1 | 2.0 | **Corporate (organisation-sponsored)** | Level | CTX | OUT | AUTH | CRED | HRS | MTH | PRC | |---|---|---|---|---|---|---|---| | L1 | 17.2 | 11.3 | 18.4 | 11.8 | 13.6 | 18.4 | 9.3 | | L2 | 17.2 | 12.3 | 17.4 | 11.8 | 15.6 | 17.4 | 8.3 | | L3 | 20.2 | 14.3 | 15.4 | 9.8 | 17.6 | 17.4 | 5.3 | | L4 | 21.2 | 16.3 | 14.4 | 8.8 | 18.6 | 16.4 | 4.3 | | L5 | 20.2 | 17.3 | 15.4 | 6.8 | 21.6 | 15.4 | 3.3 | | L6 | 21.2 | 15.3 | 16.4 | 6.8 | 22.6 | 15.4 | 2.3 | Every row sums to exactly 100.0 (largest-remainder rounding at one decimal place). Read the two tables side by side and the structural difference in the market is visible: **a personal buyer is buying a relationship; an organisation is buying a process.** Chemistry and authenticity take 30–34% of a personal buyer's weight and 14–18% of a corporate one. Method transparency and credential — the things procurement can put in a file — take 22–30% of a corporate buyer's weight and 12–19% of a personal one. Neither buyer is wrong. They are buying different goods, and a single national ranking cannot serve both. ### 5.5 The weights against the global literature The weights were derived from what buyers said, not from what the research says. They were then checked against the research — separately, afterwards, and without adjusting them to fit. That order matters: a model tuned to agree with the literature would tell you only that its author had read the literature. | Pillar | Weight (corporate L5) | What the international evidence says | Verdict | |---|---|---|---| | **Hours & Practice Depth** | 21.6% | Coach experience and contextual understanding are repeatedly identified as predictors of outcome, and the strongest reported predictor is time spent on similar problems in similar organisations. | **Validated** | | **Context Fit** | 20.2% | Coach knowledge of the client's business is among the factors sponsors tie to success (32.4% in the ICF study); contextual fit is emphasised across the effectiveness literature over status signals. | **Validated** | | **Outcome Evidence** | 17.3% | de Haan & Nilsson's RCT-only meta-analysis — the strictest filter yet applied to this field — finds effects "significant and moderate." Honest researchers report effect sizes; salespeople report ROI percentages. A coach who measures is a coach whose claims can be checked. | **Validated** | | **Authenticity & Chemistry** | 15.4% | **Contested — see below.** | **Partly contradicted** | | **Explicit Method** | 15.4% | Baron & Morin (2009): effectiveness depends on mutual agreement on goals, on the path to them, and on interpersonal comfort. Goal contracting is a method artefact, and MTH2 measures it directly. | **Validated** | | **Credential & Ethics Floor** | 6.8% | Certification confirms training completed and is a weak predictor of client outcome; meta-analytic reviews find that primary studies frequently do not even record whether coaches held credentials. This is precisely why CSI-IN makes ethics a **gate** (G1) and credential a low-weighted score. | **Validated** | | **Straight Pricing** | 3.3% | No outcome literature. Included because buyers cite fees (21.1% sponsors, 20.4% coachees) and because commercial terms are checkable. Weighted accordingly — low. | **Buyer-derived only** | #### The chemistry problem, stated against our own interest Graßmann, Schölmerich & Schermuly (2020) meta-analysed 27 samples, N = 3,563 coaching processes, and found the working alliance related to client outcomes at **r = .41** — the strongest single finding in the coaching evidence base. That looks like powerful support for weighting Authenticity & Chemistry heavily, especially for self-funded buyers, where it takes 30–34%. It is not, quite. Boyce et al. (2010) found **coach–coachee match to be unrelated to coaching outcomes**, and the alliance literature locates the effect in the relationship that *develops*, not in the match observed beforehand. The alliance is the thing that works. The chemistry meeting is not the alliance; it is a thirty-minute sample of two strangers being agreeable. So buyers are weighting a proxy whose predictive validity is weak, for a mechanism whose predictive validity is the strongest in the field. CSI-IN does not correct this, because correcting it would mean overruling the buyers whose stated preferences are the entire basis of the weights. What it does instead is three things, all disclosed: 1. **AUTH1 measures the process, not the feeling** — whether the coach offers a structured session of 45 minutes or more against a written brief, and whether alternates are offered. That is a measure of how deliberately a coach builds an alliance, which is closer to what the evidence supports than a gut reaction is. 2. **AUTH2, the gut reaction, belongs to the buyer alone** and is one indicator of twenty-one. 3. **This paragraph exists.** A buyer who reads it may move the Authenticity slider down and the Outcome Evidence slider up, and the ranking will rebuild. That is what the sliders are for. The instrument's job is to make the trade-off visible, not to make it for you. #### What this section is not It is not a claim that CSI-IN predicts coaching outcomes. It does not, and §13.7 says so. The literature above establishes something narrower and more useful: that the dimensions buyers say they select on are, with one documented exception, the dimensions the research associates with coaching that works. Buyers are mostly asking the right questions. They have had no way to check the answers. ### 5.4 Declared exclusions | Excluded | Survey value | Why | |---|---|---| | Visibility | 2.6% (E5) | Measures marketing spend, not coaching capability. Including it lets reach buy rank — the failure mode of every media list in this category. | | Coachee readiness | 57.4% (E17) | An attribute of the buyer. Moved to the Engagement Readiness Check. | | Organisational culture | 35.1% (E22) | Sponsor-side. Not coach-attributable. | | Manager support | 21.6% (E22) | Sponsor-side. Not coach-attributable. | Excluding a factor that scored 57.4% requires saying so loudly, which is the purpose of this table. Nothing is dropped silently. --- ## 6. Evidence tiers — the anti-puffery multiplier Every indicator carries an evidence tier alongside its 0–4 code. The tier multiplies the score. | Tier | Multiplier | Definition | |---|---|---| | **T1 Verified** | 1.00 | A document, register entry or record the buyer can open. | | **T2 Corroborated** | 0.90 | A named, contactable third party, or a dated public artefact. | | **T3 Declared** | 0.70 | The coach's own unverified claim. | | **T4 Absent** | 0.00 | No evidence offered. Scores zero however the claim is worded. | This is the mechanism that makes the index hard to game. A coach claiming 5,000 hours with nothing to show scores **T3 × code 4 = 2.8 of 4**, not 4. A coach claiming 5,000 hours with a credential-body log scores the full 4. The difference between an assertion and a receipt is worth 30% of every point in the model, and the buyer can see exactly where it was applied. --- ## 7. Eligibility gates — screening before scoring Weighted sums are compensatory: enough strength anywhere offsets weakness everywhere. For a purchase this consequential that is unacceptable, so five non-compensatory gates run first. | Gate | Rule | Consequence | |---|---|---| | **G1** | Written confidentiality undertaking (CRED3 ≥ 2) | **Ineligible.** Not scored, not ranked, at any level. | | **G2** | Logged hours floor: 100+ at L1–L2, 500+ at L3–L4, 1,500+ at L5–L6 | **Ineligible** at that level. | | **G3** | At least one documented engagement at or above the buyer's level (L4+) | **Ineligible** at that level. | | **G4** | A real chemistry session is offered (AUTH1 ≥ 2) | **CSI capped at 70.** Cannot reach Band A. | | **G5** | Corporate buyers: some written outcome measure (OUT1 ≥ 1) | **CSI capped at 70.** | A coach failing G2 at L5 may be entirely eligible at L2. Gates are level-specific, not verdicts on the coach. --- ## 8. The score, the uncertainty, and why we publish bands ### 8.1 Arithmetic For pillar *p* with indicators *i*, code *x* ∈ {0..4} and tier multiplier *m*: ``` PillarScore_p = 100 × ( Σᵢ (xᵢ / 4) × mᵢ ) / 3 CSI = Σₚ Wₚ × PillarScore_p / 100 → 0 … 100 ``` That is the whole model. It is a weighted additive value function over evidence-discounted, absolutely-anchored ordinal codes. It fits on one line because it has to: a buyer who cannot recompute their own result by hand does not have a transparent instrument. ### 8.2 Coverage and the confidence interval Let **κ** be the weighted share of indicators carrying T1 or T2 evidence. ``` half-width = (1 − κ) × 12 CSI reported as CSI ± half-width ``` A coach documented at every point (κ = 1) gets a point estimate. A coach who is mostly assertion (κ = 0.3) gets ±8.4 — and the buyer sees immediately that the number is soft. ### 8.3 Bands, not spurious ranks | Band | CSI | |---|---| | A | 80.0 + | | B | 70.0 – 79.9 | | C | 60.0 – 69.9 | | D | 50.0 – 59.9 | | E | below 50.0 | **Where confidence intervals overlap, coaches share a rank position and are listed alphabetically.** CSI-IN will not tell a buyer that a coach on 74.2 ± 7 beats one on 73.8 ± 7. That distinction does not exist in the data, and manufacturing it would be the same false precision we criticise in §1. Ties that must still be broken are broken deterministically and in this order: higher κ, then Outcome Evidence, then Hours, then alphabetical by registered name. There is no randomness anywhere in CSI-IN. --- ## 9. Reproducing the same result on any AI platform This is the hardest requirement in the brief, and it deserves a straight answer rather than a marketing one. **You cannot get identical rankings by asking different AI models to judge coaches.** The published evidence on this is unambiguous: within a single model, rubric-coding agreement typically sits above α ≈ 0.80, but *across* model families it falls to roughly **α ≈ 0.55–0.61 — "poor agreement"** on the standard Krippendorff thresholds, where ≥ 0.80 is reliable and < 0.67 is unreliable. Any vendor claiming their AI ranking is identical across platforms is either not testing it or not telling you. CSI-IN reaches reproducibility by a different route: **it removes judgement from the scoring path entirely.** Six mechanisms do the work. ### 9.1 No relative normalisation The largest hidden source of divergence in MCDA-style rankings is set-dependence. TOPSIS, entropy weighting, min–max and z-scoring all compute a candidate's score *relative to the other candidates*, so adding or dropping one coach silently moves everybody. Two platforms that retrieve slightly different candidate sets then produce different rankings from identical judgement. CSI-IN therefore rejects the standard MCDA toolkit and scores every indicator against a **fixed external anchor**. A coach's 78.1 is 78.1 whether they are compared with two rivals or two hundred, this year or next. This is also why a coach can be told their score and improve it — a relative index cannot offer that. ### 9.2 A frozen candidate register Models never *recall* who the coaches are. Candidates come from a versioned, hash-stamped register file, or from the buyer's own shortlist. Free recall is the single biggest source of cross-platform variance and it is designed out. ### 9.3 A codebook of countable tests Every anchor is a count or a document check. "Documented engagements in the buyer's sector in the last 5 years: none / 1 / 2–4 / 5–9 / 10+" leaves very little room for a model to differ from another model. "Rate their industry knowledge" leaves all of it. ### 9.4 Evidence-locked extraction Every point awarded must cite a field in the evidence pack. A model may not use outside knowledge, inference, or reputation. No citation, T4, zero. ### 9.5 Arithmetic outside the model The language model emits **only** the 21 integer codes and 21 tier labels — a JSON array. It never multiplies, never weights, never ranks. That work is done by `csi_reference.py`, or the JavaScript in the calculator, or a spreadsheet. Given identical codes, every platform necessarily produces an identical number, because no platform is doing the sum. ### 9.6 Decoding discipline and the adjudication rule Temperature 0, top-p 1, fixed prompt, strict JSON with fixed key order. And for the residue that remains: run the extraction on two or more platforms, and **where they disagree on a code, take the lower one and log the disagreement.** This makes the output deterministic even under disagreement, and it errs toward the buyer rather than the coach. ### 9.7 What we therefore claim, and what we do not **We claim:** given an identical evidence pack and codes, every platform and every implementation produces a bit-identical score. This is guaranteed by construction, and the conformance fixture in the spec file lets anyone verify it in one command. **We do not claim:** that two platforms reading the same evidence pack will always emit identical codes. They will mostly agree, because the codebook is countable — but "mostly" is an empirical quantity, not a promise. So we measure it. Each release publishes a **reproducibility audit**: the extraction run three times on each of three platforms, reporting Krippendorff's α on indicator codes and Kendall's τ on final ranks. Indicators falling below **α = 0.80** are rewritten or removed before the version ships. An instrument that publishes its own disagreement rate is the only kind a buyer should trust. --- ## 10. Sensitivity analysis A transparent model must show how much its answer depends on its choices. **Weight sensitivity to Δ (corporate, L5):** | Δ pp | CTX | OUT | AUTH | CRED | HRS | MTH | PRC | |---|---|---|---|---|---|---|---| | +0 | 20.2 | 11.3 | 15.4 | 6.8 | 27.6 | 15.4 | 3.3 | | +4 | 20.2 | 15.3 | 15.4 | 6.8 | 23.6 | 15.4 | 3.3 | | **+6** | **20.2** | **17.3** | **15.4** | **6.8** | **21.6** | **15.4** | **3.3** | | +8 | 20.2 | 19.3 | 15.4 | 6.8 | 19.6 | 15.4 | 3.3 | | +12 | 20.2 | 23.3 | 15.4 | 6.8 | 15.6 | 15.4 | 3.3 | **Rank stability:** across the full range Δ = 0…12, the worked example's ordering does not change. Run `python3 csi_reference.py --sens` to reproduce. Buyers who disagree with Δ can set it to zero and see for themselves. --- ## 11. Worked example Three coaches, same evidence, two different buyers. **Corporate sponsor hiring for a CXO (L5):** | # | Coach | CSI | Confidence | Band | κ | |---|---|---|---|---|---| | 1 | A — evidence-led practitioner | 78.1 | 73.8 – 82.4 | B | 0.64 | | 2 | B — senior but lightly documented | 57.8 | 49.8 – 65.8 | D | 0.33 | | — | C — credentialled, little else verifiable | — | — | **INELIGIBLE** (G1, G2) | — | **Personal buyer, first-time manager (L1):** A 77.3 (B) · B 60.2 (C). Three things this example is designed to show. Coach B is *more senior* than Coach A and still loses by twenty points — because seniority that cannot be evidenced is discounted to 70% and the thin κ widens the interval until the buyer can see the softness. Coach C holds an MCC and is **ineligible**, because a credential is a floor, not a score. And the gap narrows sharply between L5 and L1 — at L1, chemistry carries 34.4% and Coach B's relational strength nearly closes the distance. --- ## 12. Governance and conflict of interest CSI-IN is published by Nirvedha Executive Coaching Solutions, whose founder is himself an executive coach in the Indian market. That is a conflict, and it is managed explicitly rather than disclosed and ignored. The management is not abstention. An instrument whose author refuses to be measured by it tells a buyer only that the author is afraid of it, and an author absent from his own table has avoided the test rather than passed it. The author is therefore scored *and* ranked like anyone else — and gives up the two things that would actually be worth gaming: he does not code his own entry, and he does not choose when to publish it. The design fact that makes this survivable is on every page of the calculator. **Any reader can re-rank the whole table against their own weights in ten seconds**, using the sliders. A ranking nobody can recompute needs its author removed from it. A ranking anyone can recompute does not — the suspicion that the weights were shaped to flatter the author is testable by the person holding the suspicion, which is the only place it can ever be settled. 1. **The author is scored and ranked, but never codes himself.** CSI-IN has no opinion about who belongs on a buyer's shortlist and no idea whose name is in the box. Sudhakar Reddy Gade is scored like anyone else and takes his place in the table like anyone else — an instrument that exempts its own author is not being transparent, it is avoiding the test. What he does **not** do is code his own entry. Nirvedha's indicators are coded and verified by the independent adjudicator (§8.2 of the Register Protocol), from the same evidence pack any other coach submits, and the published entry is marked **codes set by independent adjudication**. The conflict was never the author's presence in the table. It was grading his own paper. 2. **The author's scorecard is published first, and published wherever it lands.** Before any other coach is scored in public, Nirvedha's completed CSI-IN scorecard is published in full — all 21 codes, every evidence tier, and the underlying evidence pack — so that any reader can recompute the number and challenge any cell. The commitment to publish is made before the result is known. An instrument that goes public only when it flatters its author is not an instrument, and a mid-table finish by the author would be the most credible thing that could happen to this one. 3. **The author codes a competitor only under the Register Protocol, never alone.** Every code traces to a citation a third party can open and check against the published anchor; there is no cell where the compiler's opinion is the input. Anything that cannot be settled by a document is left unrated rather than judged. Any cell a coach disputes is re-coded by the independent adjudicator. See `REGISTER-PROTOCOL.md` §8. 4. **Hosted independently.** CSI-IN is deployed on its own domain, separate from the author's business site, with no shared navigation, branding, analytics or lead capture. The page runs entirely in the reader's browser and transmits nothing. An instrument that lives inside a competing coach's marketing funnel invites exactly the suspicion this design exists to remove — so it does not live there. 5. **No commercial relationship with ranked coaches.** No listing fees, no sponsorship, no advertising, no referral commission, no paid removal. 6. **Right of reply.** Any coach may inspect their own indicator codes and evidence citations, submit corrections with documentation, and have a correction published in the changelog. 7. **Open by default.** Specification, weights, code and conformance fixtures are CC BY 4.0. Fork it, criticise it, publish a better one. 8. **Version freeze.** Weights are frozen per version. Changes ship as a new version with a changelog and a re-run of the previous edition under the new weights, so movement caused by methodology is never mistaken for movement caused by coaches. 9. **Coaches may be sold help, never position.** Nirvedha sells consulting to coaches on implementing the disclosure checklist. That is the rating-agency conflict in miniature — the party setting the criteria also selling the service of meeting them — and it is contained by four rules, not by good intentions. 1. **The method is published free and complete** (`FOR-COACHES.md`), generated from this same specification. There is no withheld tier and no private advice. A coach who follows it alone scores exactly what a coach who paid would score, because the score is a function of evidence, not of who assembled it. 2. **Paying cannot move a code.** Any coach who is, or was within 24 months, a paying client of Nirvedha has their register entry coded by the **independent adjudicator**, never by the compiler. 3. **The relationship is published on the entry.** Every register row carries `commercial_relationship`, and it is displayed, not buried in a footnote. 4. **No fee, ever, for listing, scoring, re-scoring, correction or removal.** Corrections are free permanently. The moment position can be bought this becomes the search engine it was built to replace. Rule 1 is the load-bearing one. Selling labour that anyone could perform themselves, using instructions anyone can read, is a service. Selling access to information that determines rank is a racket. The difference is whether the instructions are public — so they are. 10. **Two operating modes.** *Shortlist mode* — the buyer scores coaches they are already considering; nothing is published. *Register mode* — a published league table drawn from a consented, evidence-verified register, with right of reply, entered only by coaches who have submitted evidence. **v1.0 ships shortlist mode**, which is a comparison instrument, not a discovery one (§13.8). Register mode is specified and will not launch until the register holds enough consented, verified entries to be fair to the coaches in it. --- ## 13. Limitations Stated plainly, because a methodology that hides these is a brochure. 1. **The sponsor sample is small.** n = 38. A multiple-response percentage from 38 people has a wide confidence interval; the gap between "references" at 21.1% and "fees" at 21.1% is noise, not signal. The weights are the best available evidence, not precise truth. 2. **Two base cells are imputed.** Personal-buyer Context Fit and Outcome Evidence are borrowed from the corporate vector because Exhibit 17 never offered those options. Both are flagged. 3. **The level tilt is judgement.** No survey segments selection criteria by leadership level. It is bounded to ±5pp and zero-sum precisely because it is not evidence. 4. **The 50/50 split of "coach expertise" is judgement.** A different split moves the personal HRS/MTH balance. 5. **Evidence tiers reward documentation, which correlates with practice size.** A superb solo coach who keeps poor records will score below their true quality. This is a real bias, and it is the price of not rewarding assertion. Mitigation: the tiers reward *any* verifiable artefact, not institutional scale. 6. **Δ and the tilt are the author's judgements.** Both are exposed as switches so a buyer who disagrees can neutralise them and re-run. 7. **Nothing here measures whether coaching worked for you.** CSI-IN ranks the *evidence a coach can produce before you hire them*. That is a selection instrument, not an outcome guarantee — and the ICF study is clear that the largest success factors are your readiness, your goals and your manager's support, none of which any coach controls. 8. **It de-biases the comparison, not the discovery.** This is the most serious limitation in the list. In shortlist mode CSI-IN scores the candidates a buyer brings it, and says nothing about the ones who never made the shortlist. Wherever those names came from — a search engine, LinkedIn, a credential directory, a consultant returning a favour — that source's bias has already done its work before the first indicator is coded. Scoring a skewed candidate set impeccably yields an impeccable ranking of a skewed set. The calculator now runs this search itself, which closes the buyer's loop but does not close this gap — it relocates it. What the search returns is **candidates, not findings**: it is not scored, not ranked, and explicitly **not reproducible**, because web retrieval varies by provider, by index and by day. That is the one part of this instrument that cannot be made deterministic, which is exactly why it is kept outside the model: candidate discovery feeds the scorecard, and the scorecard alone produces a number. Two things follow. Buyers should assemble the widest list they can stand before scoring, because the instrument cannot reward a coach it is never shown. And register mode (§12.10) exists precisely to close this gap: a consented, evidence-verified register is the only way an instrument like this reaches coaches who are excellent and invisible — which, given that sponsors rank visibility at 2.6% while every discovery channel ranks it first, is the whole population the market is currently mispricing. --- ## 14. Files | File | What it is | |---|---| | `METHODOLOGY.md` | This document. | | `csi-spec-v1.0.json` | Machine-readable specification: pillars, anchors, tiers, gates, all 12 weight vectors, conformance fixture, SHA-256. | | `csi_reference.py` | Reference implementation. Derives the weights from the raw survey data, runs the invariants and the sensitivity analysis. No dependencies. | | `PORTABLE-PROMPT.md` | The paste-into-any-AI extraction protocol. | | `index.html` | The buyer's calculator. Runs entirely in the browser; nothing is transmitted. | | `build_html.py`, `gen_prompt.py` | Regenerate the prompt and re-inline the spec after any change to the reference implementation. | | `FOR-COACHES.md` | The coach-facing companion: how to become discoverable without spend, generated from the same spec buyers score against. | | `REGISTER-PROTOCOL.md` | How a published league table of named coaches is compiled, cited, noticed and corrected. | | `register-schema.json`, `register-sample.json` | The register file format and a worked example. | | `README.md`, `netlify.toml` | Standalone deploy. CSI-IN is hosted independently of its author's business site. | ``` python3 csi_reference.py # tables, worked example, invariants, sensitivity python3 csi_reference.py --json # regenerate the spec ``` --- ## 15. References - Kirtane, J., Kaur, R., Saksena, G., Bhowmick, A., et al. (2026). *Coaching Forward: State of Executive Coaching in South and Southeast Asia — Trends 2025–26.* ICF chapters of South and Southeast Asia. n = 397. Ethical approval: Henley Business School, University of Reading. - Dawes, R. M. (1979). The robust beauty of improper linear models in decision making. *American Psychologist*, 34(7), 571–582. - de Haan, E., & Nilsson, V. O. Randomised controlled trials in coaching — the strictest filter yet applied to the coaching evidence base; effects "significant and moderate." - Vistage Research Center (2026). *The Top Executive Coaching Companies of 2026* — weighted factor methodology. https://vistage.com/research-center/member-experience/top-executive-coaching-companies/ - Thinkers50 (2026). *Ranking methodology — viability and visibility across ten criteria.* https://thinkers50.com/t50-ranking/ - Krippendorff, K. *Content Analysis: An Introduction to Its Methodology* — α thresholds: ≥ 0.80 reliable, 0.67–0.79 tentative, < 0.67 unreliable. - Reliability-without-validity literature on LLM-as-judge agreement across model families (α ≈ 0.55–0.61). arXiv:2606.19544; arXiv:2510.27106. - Graßmann, C., Schölmerich, F., & Schermuly, C. C. (2020). The relationship between working alliance and client outcomes in coaching: A meta-analysis. *Human Relations*, 73(1), 35–58. 27 samples, N = 3,563; r = .41. - Boyce, L. A., Jackson, R. J., & Neal, L. J. (2010). Building successful leadership coaching relationships: Examining impact of matching criteria in a leadership coaching program. *Journal of Management Development*, 29(10), 914–931. - Baron, L., & Morin, L. (2009). The coach–coachee relationship in executive coaching: A field study. *Human Resource Development Quarterly*, 20(1), 85–106. - Jones, R. J., Woods, S. A., & Guillaume, Y. R. F. (2015). The effectiveness of workplace coaching: A meta-analysis of learning and performance outcomes from coaching. *Journal of Occupational and Organizational Psychology*, 89(2), 249–277. - Solms, L., et al. (2025). It's a match! The role of coach–coachee fit for working alliance and effectiveness of coaching. *Journal of Occupational and Organizational Psychology*. - Athanasopoulou, A., & Dopson, S. On the limits of self-report and the scarcity of recorded credential data in coaching outcome studies. - Hwang, C.-L., & Yoon, K. TOPSIS and the set-dependence problem in multi-criteria decision analysis — the reason CSI-IN uses absolute anchoring instead. --- *CSI-IN is published as a public good. If you can build a better instrument, please do, and tell us where ours is wrong.*