Home › Topics › Choosing a Server › Deciding

Scoring the Options and Making the Decision

The trial is over, the results sheets are full, the vendors have answered their questions, and somebody has booked a room. What usually happens in the room: the person with the strongest opinion talks the longest. A spreadsheet appears that happens to agree with them. A decision is announced that half the table privately doubts. Six months later, when the server misbehaves, the doubters say "we said so," and nobody can reconstruct why the choice was made.

A decision record is the alternative. It has scores that trace to trial evidence and bands agreed before anyone scored. It has a memo that says why this and on what grounds, written for someone who was not in the room. This article builds that record. It covers the scoring worksheet, the bands, and three worked outcomes that show the arithmetic behaving (one of them eliminates every candidate). It also covers reference checks, cost shape, and the memo. It is part of our Choosing a File Transfer Server series. It assumes the criteria matrix with its weights, the trial results sheet, and the vendor answers are in hand.

What a Score Is For

First, the disclosure the series repeats: Sysax sells a file transfer server, so this worksheet is one we would be scored on. It is built to be able to eliminate us; the second worked outcome below does exactly that. If the method cannot produce a result a vendor would dislike, it has been rigged, and you should fix it before trusting it.

A score does not make the decision. People make the decision; the score makes the reasoning explicit enough to be checked. That matters because a weighted score is only as honest as its inputs. Every input, the weights, the anchors, the evidence, was chosen by someone. The worksheet makes each of those choices visible, dated, and attributable. So when the result surprises someone, the conversation is about a specific row rather than the whole exercise.

Recall the shape from the criteria article: gates are pass or fail and are applied first. Only survivors are scored. Could rows never enter the score and only break ties. Everything below lives inside that shape.

The Scoring Worksheet

The worksheet is the criteria matrix with three columns added: the score, the evidence reference, and the scorer's initials. It is filled in twice, independently, by two people who both attended the trial. The sheets are reconciled row by row in a meeting. Every difference of more than one point must be explained from the evidence.

SCORING WORKSHEET — one per candidate, per scorer

SCALE (weighted rows only; gates are PASS/FAIL, checked first)
  0  absent, or verification failed
  1  present but failed part of the verification, or needed vendor help
  2  verified with caveats recorded on the results sheet
  3  verified as required
  4  verified with no caveats AND better than required in a way we will use

RULES
  - a score above 2 requires an evidence reference (results-sheet row,
    export file, timing entry, or vendor answer by question number)
  - a row with no trial observation scores 0, not "probably 3"
  - could rows are not on this sheet
  - weights were fixed on [date] by [two names]; they do not change now
  - two scorers, independent, then reconcile; differences over one point
    are argued from evidence, never averaged

ARITHMETIC
  category score  = average of its rows (0-4)
  category points = category score / 4 x category weight
  total           = sum of category points (out of 100)

CANDIDATE: ____________   SCORER: ____   DATE: ________
CATEGORY (weight)   ROW   SCORE   EVIDENCE REF          NOTES
Protocols (15)      P1    _       ____________          ______
                    P3    _       ____________          ______
Authentication (15) A3    _       ____________          ______
Security (15)       S2    _       ____________          ______
                    S3    _       ____________          ______
Logging (10)        L2    _       ____________          ______
                    L3    _       ____________          ______
Automation (5)      U1    _       ____________          ______
                    U2    _       ____________          ______
Administration (20) D1    _       ____________          ______
                    D2    _       ____________          ______
                    D3    _       ____________          ______
Performance (5)     F1    _       ____________          ______
Licensing (10)      C1    _       ____________          ______
Support (5)         T1    _       ____________          ______
                    T2    _       ____________          ______
                                            TOTAL (of 100): ____

Two of the rules do most of the work. The evidence rule, no score above two without a citation, stops the sheet from recording impressions, because an impression has nothing to cite. The independence rule stops one persuasive person from scoring for the room. Together they make a low score for the favorite something that has to be argued away with evidence, rather than something that quietly does not happen.

Scoring Bands

A total out of one hundred needs translating into a recommendation, and the translation is best agreed before any candidate is scored. The bands below work for most teams; the exact thresholds matter less than having written them down in advance. The diagram shows them, with gate failures outside the scale entirely.

Scoring bands on a scale of zero to one hundred. Below fifty-five: do not buy. Fifty-five to seventy: viable with conditions written into the memo. Above seventy: recommend. Separately, a gate failure means no score at all. Candidates within five points of each other are treated as tied and separated by tie-breakers.

The "viable with conditions" band is the useful one. A candidate scoring in it is not rejected. It is bought only if the conditions that pulled its score down are fixed in writing. Examples include an edition confirmed, an exit term added to the contract, a support tier upgraded. The memo names each condition. A candidate in the recommend band still faces the reference checks and the cost-shape review, which can move it down, never up.

The tie rule deserves emphasis. Weighted scores carry false precision; a difference of two or three points is well within the noise of two scorers' judgment. So candidates within five points are treated as tied and separated by tie-breakers. Those are the could rows, the reference checks, the cost shape, and the support experience during the trial. Deciding a tie by the decimal is how the spreadsheet ends up making a decision it is not qualified to make.

Three Worked Outcomes

The arithmetic is only trustworthy once you have watched it behave. Three synthetic evaluations, each using a weighting column from the criteria article:

Outcome A: a clear winner, with conditions

The regional distributor from the requirements article (nine inbound partners, three outbound, two generalist admins, a Windows-only estate) used the small-partner-exchange weights. Two candidates survived the gates: a self-hosted Windows server with a console-driven admin model, and a cross-platform commercial server built around a web portal. Reconciled totals: 74 and 63.

The first candidate scored high on administration (every timed task under target, restore in twenty minutes) and authentication (directory login and key rotation verified). It scored low on automation. The arrival trigger the order-system team wanted sits in a higher edition. So the row was scored on the edition the team could budget for. The second scored high on automation and the browser upload page. It scored low on administration (the restore took over an hour and needed the vendor). It also scored low on support (the trial ticket was routed twice before reaching an engineer). Both cleared fifty-five; one cleared seventy; the gap exceeded five points.

The memo recommended the first candidate with two conditions. The first was written confirmation of which edition contains the trigger, priced before signature. The second was a contractual exit clause matching vendor answer X1. The dissent paragraph recorded the order-system team's preference for the second candidate's automation. The ninety-day review was tasked with checking whether the trigger had actually been deployed.

Outcome B: everyone fails a gate

Acme was a regulated financial services firm with a Linux-only estate and two EDI partners. It had three gates that mattered: runs on the team's Linux platform, speaks AS2, and forwards logs to the central platform with every field intact. Three candidates were on the shortlist. One was a Windows-native server (Sysax Multi Server, for the sake of the example). Another was a cross-platform commercial server. The third was a hosted suite offering the full managed-transfer layers.

The Windows server failed the first gate before any trial was installed, which is the requirements worksheet doing its job. It was never scored. The cross-platform server passed the platform gate and failed the AS2 gate in the first week of the trial. The hosted suite passed both, then failed the logging gate on day eleven. Its central-forwarding option dropped the source address from every event, and the control matrix required it. Three candidates, zero survivors, no scores. The room was quiet for a moment.

That is not a failed evaluation; it is a successful one that found the shortlist was wrong. Acme checked whether any gate was really a should. The AS2 gate was contractual, the logging gate was in the control matrix, the platform gate was staffing. Acme concluded all three stood. The memo recorded the eliminations with their evidence and started a second shortlist built from the AS2 requirement outward. Had the team skipped the gates and gone straight to weighted scoring, the hosted suite would have "won" with a high total and a log that could not satisfy the auditor.

Outcome C: a near tie, settled by tie-breakers

An internal automation hub, where most traffic is applications exchanging files on schedules, used the automation-heavy weights. Two candidates survived: totals of 71 and 69, a tie under the five-point rule. The tie-breakers decided it. The could rows favored the first slightly. The reference checks did not. The first candidate's reference administrator described an upgrade that had silently dropped the folder rules once, recovered from backup. The second's references were uneventful in the way that matters. The cost shape favored the second. The first was licensed per concurrent connection, a metric that grows with every new application. The second was one-time with maintenance. The support tickets were comparable.

The memo recommended the second candidate. The memo's dissent paragraph, written by the advocate of the first, recorded the higher raw score and the automation rows behind it. That paragraph is not a formality; if the second candidate disappoints, it is the first thing the ninety-day review will reread.

Remember: outcome B is the one that proves the method is honest. If your process has never eliminated every candidate, or never eliminated the one the room liked, look hard at whether the gates are real.

Reference Checks That Ask the Right People

A reference check is a conversation with an existing customer the vendor introduces, which means the vendor chose them. (It is still worth an hour. Nobody runs a product for years without learning something worth repeating.) The check works provided you speak with the administrator who runs the product rather than the manager who bought it. The environment you discuss must also resemble yours in operating system, size, and partner count. Then ask questions a satisfied customer can still answer unflatteringly:

  • What broke in the first year, and how did you find out?
  • Describe your last upgrade: how long it took, what changed, what you had to fix afterwards.
  • When did you last open a support ticket outside business hours, and what happened?
  • What does the product not do that you wish it did?
  • Knowing what you know now, what would you have done differently at purchase?
  • How many people can administer it at your organization, and how were they trained?

Record the answers by role, not by name. Treat a reference who cannot name anything that ever went wrong as one who has not run the product very long. Reference findings never raise a score; they can lower a category, add a condition, or settle a tie.

Cost Shape Without Spreadsheet Theater

Spreadsheet theater is a five-year total-cost model calculated to the last unit. It is built on guesses about user growth, hardware, and hours, presented as if its precision were real. It persuades executives and it is almost always wrong, because the guesses dominate the arithmetic. I have built one. It had a tab called "assumptions" that nobody opened, which was the most accurate thing in it. The alternative is to describe each candidate's cost shape in words and compare shapes over three to five years:

  • The purchase model: one-time license plus annual maintenance, or subscription. Which does finance prefer, and which did the requirements record?
  • What the cost grows with: nothing, servers, cores, users, connections, partners. A metric tied to something the business will grow (partners, applications) is a rising shape; a per-server license is flat until the next server.
  • Step functions: the edition boundary where a needed feature lives, the user count where a tier changes. Note where each step sits relative to your expected growth.
  • The costs no invoice shows: these include administrator hours, estimated from the trial timings multiplied by how often each task recurs. Other costs are training, hardware, and partner coordination at cutover. That coordination is estimated from the partner window.

One candidate might be described as "flat, one step at the edition boundary we already need, low admin hours". Another might be "rising with every new application, no steps, higher admin hours from the restore task". People who would never trust each other's spreadsheets can compare those two candidates honestly. The memo records the shapes and lets the reader do the arithmetic for their own growth assumptions. The one-page business case uses the same shape-in-hours language when the money has to be asked for. Where a hosted candidate is in the mix, the self-hosted versus hosted comparison lists the shape differences invoices hide. For whether buying anything beats building on scripts, the tipping-point worksheet handles that accounting, and it sometimes says keep building.

The Loudest-Stakeholder Trap

Every selection has one participant whose preference is stronger, louder, or more senior than the evidence. The trap is not that they are wrong; sometimes they are right. The trap is that their preference substitutes for the process, and the decision inherits none of its defensibility. Four defenses, all already built into the method:

  • Weights fixed before the trial, with two names on them, so that the stakeholder cannot re-weight the sheet after seeing the scores.
  • Independent scoring, so that the loudest voice fills in one sheet, not everyone's.
  • The evidence rule, so that "I just think it's better" cannot produce a four in any row.
  • The dissent paragraph, written by the advocate of the losing candidate and included in the memo unedited. It gives the stakeholder a formal place for their view, which is often all they wanted, and it gives the ninety-day review something to check.

One more habit helps. Before the memo is signed, spend ten minutes on a pre-mortem. Imagine it is a year from now and the choice has clearly failed. Have everyone write one sentence on why. The sentences usually point at a row that was scored generously, and it is cheaper to re-examine it now than to be right about it later.

The Decision Memo

The memo is the artifact that outlives the meeting. It is two or three pages, written for someone who was not in the room. That might be a new manager, an auditor, the administrator who inherits the server in three years. The memo answers one question: why this, on what evidence. The skeleton:

DECISION MEMO — transfer server selection          Date: ______

1. DECISION       One sentence: what we are buying, which edition, and why.
2. REQUIREMENTS   Reference to the frozen must list: date, signers, and
                  any changes since the freeze with their written reasons.
3. GATES          Each candidate, each gate, pass/fail, evidence reference.
                  Candidates eliminated before trial and at which row.
4. SCORES         Weights (date, two names). Reconciled totals per
                  candidate. Rows where scorers differed and how resolved.
5. TRIAL EVIDENCE Results-sheet reference; injections that produced
                  concerns; task timings; partner-window notes.
6. VENDOR ANSWERS Answers cited by question number that affected a score
                  or created a condition.
7. REFERENCES     Roles spoken to, environment similarity, what they said
                  that lowered a score or settled a tie.
8. COST SHAPE     Each candidate's shape in words over the horizon:
                  model, growth driver, step functions, unbilled costs.
9. CONDITIONS     What must be true in the contract before signature.
10. DISSENT       The strongest case for the runner-up, written by its
                  advocate, unedited.
11. REVIEW        Ninety-day review date, attendees, and the three
                  measures that will show whether the decision held.

Signed: administrator ____  security ____  budget owner ____

Section ten is the one teams are tempted to cut. Keep it. A memo that records only the winning argument reads, a year later, like marketing. A memo that records the best case against its own decision reads like a group of people who thought carefully. Section eleven is the bridge to the last article in this series, the ninety days after purchase. That is where the memo's measures are checked against reality. The article presenting the case and following through covers reporting them back to whoever signed the budget.

Gotcha: a memo written after the purchase order is signed is a justification, not a decision record. Write it, circulate it, and collect the signatures before anything is bought. The memo is what the signatures are on.

The Decision, Made

You now have a recommendation with a paper trail. The trail includes gates applied and recorded, and two independent score sheets reconciled from evidence. It includes a band that says what the number means, and tie-breakers where the numbers were too close to trust. It also includes references that asked the right people, cost shapes in words, and a memo with the dissent preserved. Any of it can be checked by someone who was not there, which is the whole point. Nobody has to doubt privately; the doubts are in section ten, signed.

The next article covers what happens after the signature. It covers the deployment baseline, hardening from day one, and documentation while it is fresh. It also covers the first flows and the review that checks whether the decision is holding. The decision may have been to keep building, or to split the estate between a standard SSH server and a commercial product. If so, our open source versus commercial series is the natural next read.

Frequently Asked Questions

What if the two scorers disagree by a lot on many rows?
That usually means the anchors were not specific enough or one scorer relied on impressions. Go back to the results sheet row by row and score only from what is written there. Large, persistent disagreement is information: the evidence for that candidate is thinner than it looked.
Can we buy a candidate that scored in the "do not buy" band?
Only by changing the weights or the requirements in writing, with the same sign-off that froze them, and recording why in the memo. The honest description is then "we decided on other grounds," and the memo should say so rather than adjusting the score.
Should the vendor see the scores?
Share the conditions, since the vendor has to satisfy them, and the gate failures, since those are facts. The weighted scores and the dissent paragraph are internal. A vendor who asks to review the weights is asking to negotiate the method.
Is the memo overkill for a small purchase?
Shrink it to a page, but keep sections one, three, nine, ten, and eleven. A short memo recording the gates, the conditions, the dissent, and the review date still lets the administrator who inherits the server understand the decision in five minutes.
What if the price changes after the memo is signed?
A changed amount changes the cost shape section, so the memo is amended and re-signed. The changed amount does not change the technical scores. If the new shape pushes the candidate into a worse band than the runner-up, the tie-breakers are re-run.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.