Home › Topics › Troubleshooting Method › Protocol

Layer Four: Protocol-Level Failures

"It must be a bug." Everything is allowed and it still does not work. The path exists, the login succeeds, and the account has every right it needs. Yet the listing hangs, or the TLS handshake fails, or the SFTP session closes the instant it opens with no explanation. These are protocol-level failures. The two sides can reach each other and trust each other, but they cannot agree on how to carry out the conversation. They are the failures that make administrators say "it must be a bug," and they almost never are. Both sides are following the rules. They have different copies.

This is the fourth layer of our Systematic Troubleshooting of Failed Transfers series. Protocol failures come in three families: FTP's second connection, TLS negotiation, and SSH/SFTP negotiation. Every one of them announces itself in verbose client output if you know how to switch it on and where to look. This article shows you how to capture that output with each common tool. It shows you how to read FTP reply codes and SFTP status codes properly, and what each family's telltale messages mean. Several of these problems have whole articles of their own elsewhere in the library. This one teaches you to recognize which problem you have, then points you to the fix. The client knew which family it was the whole time. Nobody had asked it to say so.

First, Get the Client to Show Its Work

Graphical clients report a summary — "Failed to retrieve directory listing" — that hides the exchange which produced it. Every serious client can print that exchange, and at this layer you should not attempt a diagnosis without it. The switches:

sftp -vvv user@host                 # OpenSSH: three v's show key exchange, algorithms, subsystem request
ssh -vvv user@host                  # same, for the SSH layer on its own
curl -v ftp://host/ --ssl-reqd      # every FTP command and reply, plus the TLS negotiation
curl -v https://host/upload/        # request and response headers, TLS details
lftp -d host                        # lftp debug mode; or "debug 3" inside a session
ftp -d host                         # Windows ftp.exe: print commands and replies
# Python ftplib:  ftp.set_debuglevel(2)
# Paramiko:       paramiko.util.log_to_file("sftp.log", level="DEBUG")
# FileZilla: raise the debug level in Settings, then read the message log
# WinSCP: Preferences > Logging > enable session logging; or /log=path /loglevel=2 on the command line

Capture the whole output to a file. Then find two lines: the last thing that worked and the first thing that did not. The command or negotiation step between them is your failure, and it will belong to one of the three families below. Reading from the top rather than the bottom matters. The final error line is usually a consequence ("Connection closed") and the cause sits several lines earlier. "Connection closed" is the log's way of saying "and then I left."

Reading Reply Codes Properly

FTP servers answer every command with a three-digit code, and each digit carries meaning. The first digit is the verdict. 1xx means "started, wait for more"; 2xx means success. 3xx means "fine so far, send the next part." 4xx means failed, but temporary — try again later. 5xx means failed, permanent — retrying without changing something is pointless. The second digit is the subject: x0x syntax, x2x connections, x3x authentication, x5x the file system. So a 425 is "connection problem, temporary" and a 553 is "file system problem, permanent" before you have read a single word of the text. The most useful codes at this layer:

Code Meaning What it usually points at
150 Opening data connection Normal. If nothing follows it, the data connection never completed.
227 / 229 Entering passive / extended passive mode Decode the address and port; a private address here is the announced-address problem.
234 Security exchange accepted (after AUTH TLS) Explicit FTPS is supported; the TLS handshake comes next.
421 Service not available, closing control connection Idle timeout, connection limit reached, or server shutting down.
425 Can't open data connection Data channel blocked: mode, passive range, or firewall.
426 Connection closed; transfer aborted Data connection dropped mid-transfer: timeout, reset, or a session-reuse requirement on FTPS.
500 / 502 Command unrecognized / not implemented The client used an optional command the server lacks: AUTH, EPSV, MLSD, OPTS.
503 Bad sequence of commands The client skipped a required step, often PBSZ/PROT before an FTPS transfer.
530 mid-session Not logged in After a successful login, usually "encryption required for this operation."
550 / 553 File unavailable / name not allowed Layer three — permissions or a missing path, not protocol.

SFTP has a shorter list. The server returns a numeric status and the client prints a phrase. The phrases are 0 OK, 1 end of file, 2 no such file, 3 permission denied, and 4 failure (the catch-all). They continue with 5 bad message, 6 no connection, 7 connection lost, and 8 operation unsupported. Codes 5 through 8 are the protocol-layer ones: a client and server disagreeing on packet format, or a session dying underneath the file operations. The full FTP vocabulary is in FTP commands and reply codes; here the skill is reading the digits before the words. The digits were standardized; the words were left to whoever was typing.

Family One: FTP's Second Connection

FTP's distinguishing feature is that commands and data travel on separate connections, and the second one is negotiated per transfer. That negotiation is the single most common protocol-layer failure in file transfer, and its signature is unmistakable. Login succeeds, then the first LIST or transfer hangs, or fails with 425 Can't open data connection. Or you get a client-side Failed to retrieve directory listing after a 150 that is never followed by 226. In verbose output, look at the reply to PASV:

> PASV
< 227 Entering Passive Mode (10,0,0,15,195,87)
> LIST
*   Trying 10.0.0.15:50007...
* connect to 10.0.0.15 port 50007 failed: Connection timed out

The server announced 10.0.0.15 — a private address that only exists inside its own network — so the client obediently connected to nowhere. There are three questions for this family. Which mode is in use? A PORT before the transfer means active; PASV or EPSV means passive. What address and port did the server announce, and is that address reachable from here? Is the announced port within a range the server's firewall actually allows? On the server side, a Windows product such as Sysax Multi Server exposes the passive port range and the announced external address as ordinary settings. The fix is nearly always one of those two plus the matching firewall rule. The full diagnosis — decoding the six numbers, switching modes as a test, active-mode callbacks blocked at the client — is in diagnosing FTP mode failures. The general firewall side is in our Firewalls, NAT, and File Transfer series.

FTPS adds a twist that catches people who have already fixed passive mode for plain FTP. Once the control connection is encrypted, a firewall's FTP helper can no longer read the PASV reply and open the data port for it. So a configuration that worked for FTP fails for FTPS with the same 425. Some servers also require the data connection to reuse the control connection's TLS session. They refuse with a 4xx or 5xx reply mentioning session reuse when an older client does not. The helper cannot read what it cannot see, and it does not mention that. Both are covered in FTPS through firewalls and NAT. What the helper is doing in the first place is explained in ALG helpers and inspection.

Family Two: TLS Negotiation Mismatches

TLS failures happen before any FTP or HTTP command, so the fingerprint is a handshake error with no reply code at all. The verbose output stops at the negotiation, and the wording tells you which of four things went wrong.

The wrong kind of FTPS. Explicit FTPS connects to port 21 in plain text and upgrades with AUTH TLS; implicit FTPS expects TLS from the first byte on port 990. Cross them and you get two very specific symptoms. A client that speaks TLS immediately to port 21 receives the server's plain-text 220 greeting where it expected a TLS handshake. It fails with wrong version number — an OpenSSL message that, despite its wording, means "that was not TLS at all." A client that speaks plain FTP to port 990 sends its USER and waits forever, because the server is waiting for a handshake that never comes. Check which the job is configured for against which the server offers; explicit vs implicit FTPS has the full comparison.

The server does not offer TLS on this port. AUTH TLS is answered with 500 Unknown command or 502 Command not implemented. Either FTPS is not enabled, or it is enabled on 990 only. Conversely, a server that requires TLS answers a plain-text login with a 5xx reply saying encryption is required — and the fix is on the client.

No common protocol version or cipher. Both sides support TLS, but not the same flavor of it. One is configured to accept only recent protocol versions and strong cipher suites. The other offers only old ones. The messages are alert handshake failure, alert protocol version or no protocols available. The quickest confirmation is openssl s_client — for explicit FTPS, openssl s_client -connect ftp.acme.example.com:21 -starttls ftp. It prints the negotiated protocol and cipher on success and the alert on failure. The right fix is to modernize the old side. The temporary fix is a client or server setting that re-enables an older cipher, documented as an exception. I have found "temporary" cipher exceptions older than the server they were on. What those settings mean is explained in cipher policy basics.

Certificate trust. The handshake reaches the certificate and the client refuses it. You may see curl's SSL certificate problem: unable to get local issuer certificate or lftp's Certificate verification: Not trusted. Or you may see a graphical client's "unknown certificate" dialogue that a scheduled job cannot click. Also a server demanding a client certificate the job does not present, which fails right after the server's certificate request. Recognize these at this layer and fix them in installing and chaining certificates. The usual cause is a missing intermediate certificate on the server, not a problem on the client at all. The client is refusing correctly, which is the least popular kind of correct.

Remember: a TLS error with the phrase wrong version number almost always means you spoke TLS to a port that expected plain text. Check explicit versus implicit before you touch a single cipher setting.

Family Three: SSH and SFTP Negotiation

SFTP runs as a subsystem inside an SSH session. The client authenticates over SSH, then asks the server to start the SFTP service on that connection. Failures here happen after a successful login and before the first file operation, and each has a distinctive line in -v output.

The subsystem is not there.

debug1: Sending subsystem: sftp
subsystem request failed on channel 0
Connection closed

SSH works — you could probably open a shell — but the server has no SFTP subsystem configured for this user, or its configured path is wrong. On the server, the log says subsystem request for sftp by user feed_acme failed, subsystem not found. The fix is the Subsystem sftp line in the server configuration, covered in SFTP server configuration.

The shell talks over the protocol.

Received message too long 1416128883
Ensure the remote shell produces no output for non-interactive sessions.

This looks like corruption and is actually a login script. SFTP packets begin with a four-byte length. Suppose the account's shell start-up file prints anything — a banner, an echo, a warning about a missing directory. The client reads the first four characters of that text as a length. The number is the giveaway: 1416128883 is the bytes T h i s, the start of a message such as "This system is for authorized users only." Silence the shell for non-interactive sessions (guard the output with a test for an interactive terminal). Or give transfer-only accounts the server's built-in SFTP handler and no shell at all.

Bluewater Bank's version of this arrived on a Tuesday, courtesy of a policy push that added a legal notice to every login shell on the estate. Every SFTP job that logged in to those hosts failed overnight with Received message too long 1416128883. The first suggestion in the incident channel was disk corruption. The number decoded to This, the first word of "This system is for authorized users only." The security team's change log had the banner going in at the exact minute the first job died. The notice stayed, because it had to; the profile now prints it only when a terminal is attached. Every job ran the next night, and the banner has been quietly not corrupting anything ever since.

No common algorithm.

Unable to negotiate with 203.0.113.10 port 22: no matching host key type found. Their offer: ssh-rsa

The same message appears for key exchange methods and ciphers — no matching key exchange method found, no matching cipher found. It is always followed by Their offer: listing what the other side supports. One side has retired an algorithm the other still insists on. The right fix is to update the old side. A stop-gap is a client option that re-enables the named algorithm for that host only. Write it into the job's configuration with a comment explaining why. Our how SSH protects transfers article explains what these algorithms do.

The session dies right after login. Suppose you get a bare Connection closed immediately after authentication, with the server log showing fatal: bad ownership or modes for chroot directory. That is the jail-ownership rule from Layer Three surfacing at this layer. A Connection closed or Couldn't read packet: Connection reset by peer some time into a session, on a large or slow transfer, is a dropped session. The cause is an idle timeout, a NAT table expiring, or a missing keepalive. That is the territory of our Timeouts, Keepalives, and Dropped Sessions series. The client-side message client_loop: send disconnect: Broken pipe belongs to the same family.

Feature Mismatches That Look Like Bugs

The last group is quieter: the connection works, the transfer mostly works, and one optional feature fails. The verbose log shows a 500 or 502 in reply to a command you did not know your client sent. I learned what EPSV was from a server that refused it.

  • Modern listing commands. Clients prefer MLSD for machine-readable listings; an old server answers 500. Most clients fall back to LIST silently; some do not, and the symptom is an empty or failed listing on one server only. Clients have a setting to disable MLSD.
  • Extended passive mode. Clients send EPSV before PASV; a server or firewall helper that does not understand it produces a hang or a 500. curl's --disable-epsv and lftp's set ftp:prefer-epsv false make the client use plain PASV. The four commands are compared in PORT, PASV, EPRT and EPSV.
  • Resume. A client that tries to continue an interrupted upload sends REST. A server that does not support it replies 502, and the client either restarts from zero or fails. Whether resume is supported on both ends is the subject of our Resume and Checkpoint Restart series.
  • Character sets. A client sends OPTS UTF8 ON; the server rejects it; file names with accented characters then arrive garbled or fail to open. That is a content problem in disguise, covered in ASCII, Binary, and Encoding Corruption.

The pattern for all of these is the same: identify the optional command from the verbose log, turn it off on the client for that one server, and record the exception. A feature mismatch is never fixed by retrying. It can, however, be retried for a very long time.

Symptom to Family to Test

You see Family One test Usual fix
Login OK, LIST hangs or 425 FTP data channel Read the 227 address and port; try the other mode Passive range + announced address + firewall rule
Works over FTP, 425 over FTPS FTP data channel (encrypted) Same test with --ssl-reqd Static passive range open on the firewall; no reliance on the helper
wrong version number, or plain FTP hangs before 220 TLS (explicit/implicit) Port 21 with -starttls ftp vs port 990 without Match the client's FTPS type to the server's port
alert handshake failure, no protocols available TLS (version/cipher) openssl s_client shows the negotiated or refused cipher Modernize the old side
unable to get local issuer certificate TLS (trust) openssl s_client verify return code Install the intermediate on the server
subsystem request failed SFTP subsystem Server log: subsystem not found Configure the Subsystem sftp line
Received message too long SFTP subsystem ssh user@host true — does anything print? Silence the shell start-up for non-interactive sessions
Unable to negotiate ... Their offer: SSH algorithms Read the offer list in the message Update the old side; documented exception if you cannot
500/502 to an optional command Feature mismatch Identify the command in verbose output Disable that feature for this server

A Worked Example

A job that uploads over explicit FTPS from a Windows server has failed every night since the partner "upgraded security." The job log says only Connection failed. Reproduced with curl.exe -v --ssl-reqd ftp://ftp.acme.example.com/, the output shows 220, AUTH TLS, 234 AUTH TLS successful — and then curl: (35) ... alert handshake failure. Family two, and specifically the third case: the FTP side is fine, the TLS upgrade is offered, the handshake itself fails. openssl s_client -starttls ftp from the same machine fails the same way; from a newer machine it succeeds and prints a modern cipher. The job server's TLS library is too old for the partner's new policy. One change — updating the client library on that server — and the job runs. Nobody touched the firewall, and nobody asked the partner to weaken anything.

Finishing the Layer

Protocol failures reward exactly the discipline the rest of this series teaches. Capture the verbose exchange, find the last success and first failure, and name the family. Change the one setting that family calls for. What they punish is the alternative — toggling passive mode, disabling certificate checks, and downgrading ciphers until something works. That leaves a job that runs on luck and a configuration nobody can explain. Record the family and the setting in the write-up. Luck is not a setting, and it does not survive the next upgrade.

When the protocol conversation completes and the job reports success — and the file is still wrong — you have reached the last layer. That is Layer Five: The Transfer Worked but the File Is Wrong. It also covers writing up the whole investigation. If the negotiation problem turned out to be a blocked port rather than a mismatch, the tests in Layer One will prove it. And for the mechanics underneath this whole layer — how SFTP actually rides inside SSH — how SFTP works is the background read. It was not a bug. Write down what it was instead.

Frequently Asked Questions

The GUI client works but the script fails against the same server. Why?
They are almost certainly not sending the same commands. The GUI may be using passive mode, explicit FTPS and a fallback from MLSD to LIST. The script may use defaults that differ on every one of those. Capture verbose output from both and compare the command sequences line by line.
Is a 4xx reply always temporary?
By the protocol's definition, yes — it means "try again later." In practice a 425 caused by a blocked passive range will happen every time until the firewall changes. Treat 4xx as "may recover on its own"; if it recurs on every attempt, diagnose it as permanent.
Can I just turn off certificate verification to make FTPS work?
You can, and you will have removed the only protection against connecting to an impostor. The failure is telling you something real, usually a missing intermediate certificate on the server, which takes ten minutes to fix properly. Disable verification only as a documented, temporary exception.
What does "Received message too long" have to do with my shell?
SFTP expects binary packets that start with a length field. If the account's login script prints text, the client reads the first four letters of that text as the length — a huge number — and gives up. Make the shell silent for non-interactive logins, or use a transfer-only account with no shell.
How do I know whether a hang after login is a protocol problem or a firewall problem?
Read the PASV reply in verbose output. If the announced address is private or the port is outside the server's configured range, it is a server setting. If both look right and the data connection still times out, the port is blocked on the path — a firewall problem wearing protocol clothing.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.