Home › Topics › Troubleshooting Method › Authentication

Layer Two: Why Authentication Fails

Permission denied. Two words and a full stop, offered by the server as though that were a complete explanation. The banner arrived, the connection is fine, and the login was refused. In one way this is the friendliest failure in file transfer, because it proves the whole network path works. In another way it is the most dangerous, because the reflex response is to reset something. Reset the wrong thing and you have turned one broken job into two, or locked an account that was about to recover on its own.

This is the second layer of our Systematic Troubleshooting of Failed Transfers series. There are six common reasons a login is rejected. They are a wrong credential, a locked account, an expired password, a key mismatch, a changed host key, and a partner-side allowlist. Each leaves a distinct trace in the client's output and the server's log. The point of this article is that you can tell them apart from those traces before anyone resets anything. You can use the OpenSSH tools common to Linux and Windows, plain FTP clients, and the Windows account checks alongside. The server knows exactly why it said no. It was never going to tell the client.

What a Rejected Login Looks Like

First, learn to recognize the layer. Authentication failures always come after a successful connection and before any file operation, and each protocol has its own vocabulary for them:

  • FTP and FTPS: a 530 reply — 530 Login incorrect, 530 Login or password incorrect!, 530 Not logged in. The exact words vary by server; the number does not.
  • SFTP and SCP (SSH): Permission denied (publickey), Permission denied (publickey,password), or interactively Permission denied, please try again. The words in brackets list the methods the server would accept — a clue we will use shortly.
  • HTTPS uploads: a 401 Unauthorized status (credentials missing or wrong) as opposed to 403 Forbidden (credentials fine, action not allowed — that is layer three).
  • Scripted clients: a library exception carrying the same text, such as AuthenticationException: Authentication failed from Paramiko, or an Access denied line in a WinSCP session log.

There is one impostor. Host key verification failed and the block-capital WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED! are not the server rejecting you — they are your client refusing to proceed. They belong at this layer because they stop the login, but the cure is entirely different, and we treat them as cause five below.

Rule Zero: Do Not Reset Anything Yet

Resetting a password or regenerating a key feels like action. Consider what it actually does. If the cause was a locked account, the reset does nothing and the account re-locks the moment the job retries with the credential you did not update. If the cause was a partner allowlist, the reset does nothing at all. If a second job shares the account — common with service accounts — you have just broken it. And in every case you have destroyed the evidence. The old credential is gone, so you can no longer test whether it was wrong. Resetting a credential is the one troubleshooting step that also erases the crime scene.

The right sequence is: read the client output, read the server log, name the cause, and only then change the one thing that cause requires. The six causes and their fingerprints follow.

Bluewater Bank had a morning that shows why the sequence matters. Every partner login to their transfer server began failing with 530 Login incorrect at the same minute, domain accounts only. The two local test accounts still worked. The fact sheet said scope: everyone, which ruled out one stale password. The server log showed the failures beginning at the exact minute the virtual machine had been moved to a new host. The new host's clock was nine minutes off. Kerberos refuses to authenticate anyone whose clock is that far from the domain controller's. One time-sync command fixed every account at once, and the passwords nobody had reset were still good. The administrator on duty had a reset script open in another window, and closed it.

The Six Causes and Their Fingerprints

Cause 1: The credential is simply wrong

The boring cause is still the most common one, but "wrong" has more shapes than a typo. A scheduled job may read its password from a file with a trailing newline. A password containing $, ! or % may be mangled by shell or batch-file quoting. The username may be right for the partner's test server and wrong for production. The vault may hold the new password while the script still holds the old one. And on some servers usernames are case-sensitive. A trailing newline has cost me more mornings than any attacker I have met.

The test is a manual login using exactly the credential the job uses — copied from the job's own configuration, not typed from memory. For SSH, add -v and watch which method is tried:

$ sftp -v feed_acme@sftp.acme.example.com
...
debug1: Authentications that can continue: publickey,password
debug1: Next authentication method: password
feed_acme@sftp.acme.example.com's password:
debug1: Authentications that can continue: publickey,password
Permission denied, please try again.

The server offered publickey,password, the client tried a password, and the server still lists both methods as available. That is SSH's way of saying "that password was wrong; try again." Compare that with the server log, which for a wrong password on a valid account reads:

Mar 14 02:10:04 sftp01 sshd[4187]: Failed password for feed_acme from 198.51.100.7 port 51234 ssh2

and for a username the server does not know at all:

Mar 14 02:10:04 sftp01 sshd[4187]: Invalid user feedacme from 198.51.100.7 port 51234
Mar 14 02:10:06 sftp01 sshd[4187]: Failed password for invalid user feedacme from 198.51.100.7 port 51234 ssh2

That Invalid user line is gold: the account name itself is wrong (here, a missing underscore), so no amount of password-guessing will help. On the FTP side, watch the two-step exchange in a verbose client. 331 after USER means the name was accepted and a password is wanted. 530 after PASS means the password was rejected. 530 straight after USER means the account does not exist or is disabled.

Cause 2: The account is locked

Most servers lock an account after a number of failed logins within a window, either until an administrator intervenes or for a cooling-off period. The fingerprint is a credential you have just verified is correct — it matches the vault, it matches the job — and still fails. Two more signs: the failure began shortly after a password rotation, and the server log shows a burst of failures before the lock, not one.

The burst is the important part, and it usually comes from the job itself. A scheduled task that retries every minute with a stale password generates dozens of failures before anyone notices. It locks the account. Then — after the administrator unlocks it — it locks the account again within minutes, because the stale password is still in the job. Find and fix the source of the failures before unlocking, or the unlock buys you nothing. Unlock first and the job, ever diligent, re-locks it before your coffee cools.

Where to look:

  • Windows domain or local accounts: net user svc_feed /domain shows Account active Locked when locked. In PowerShell, use Get-ADUser svc_feed -Properties LockedOut. The domain controller's security log records the lockout event with the name of the machine the bad attempts came from. That identifies the job that needs its password updated.
  • Linux with PAM lockout: the authentication log shows the lock decision. Depending on the module in use, faillock --user feed_acme or pam_tally2 --user feed_acme shows the failure count and resets it.
  • Transfer servers with their own lockout or auto-blocking: the activity log records the failed attempts and the block, with the source address. Blocks are often by address rather than account. So a working credential from a blocked address fails while the same credential from elsewhere succeeds — a useful test.

Lockout design — how many attempts, how long, and how to avoid locking out your own jobs — is covered in authentication failures and lockouts. Here the task is just to recognize the signature and fix the source of the failures first.

Cause 3: The password or account has expired

Expiry is lockout's quieter cousin: nothing failed, time ran out. Password-aging policies expire service-account passwords that nobody thought to rotate. Accounts created for a project are given an end date and then the project overruns. Partner accounts are provisioned with a one-year lifetime. The fingerprint is a failure with no preceding burst, on a credential that is correct, often on the first run after a specific date. Expiry is the only failure that was scheduled in advance and still surprised everyone.

Interactive tools usually say so plainly — SSH prints WARNING: Your password has expired and demands a new one. But a scheduled job cannot answer that prompt, so it fails. The server-side line for SSH with PAM:

Mar 14 02:10:05 sftp01 sshd[4187]: pam_unix(sshd:account): expired password for user feed_acme (password aged)

Check expiry directly rather than inferring it. On Linux, chage -l feed_acme lists the password and account expiry dates — compare them with today. On Windows, net user svc_feed /domain prints Password expires and Account expires lines. Get-ADUser svc_feed -Properties PasswordExpired, AccountExpirationDate gives the same in PowerShell. On a transfer server with its own user database, the account's settings page or its activity log will show the expiry. The fix is a policy conversation, not just a reset. Service accounts need either non-expiring passwords with a documented rotation process, or key-based authentication. In either case the expiry dates belong on a watch list rather than in someone's memory (see certificate and key expiry watch). Our guide to service account hygiene covers both.

Cause 4: The key does not match

Key-based SSH authentication fails with Permission denied (publickey), and that single message hides at least five different problems. (SSH is economical with its disappointments.) The client's -v output separates them. A healthy key exchange looks like this:

debug1: Offering public key: /home/svc_feed/.ssh/id_ed25519 ED25519 SHA256:Q8xk...
debug1: Server accepts key: /home/svc_feed/.ssh/id_ed25519 ED25519 SHA256:Q8xk...
Authenticated to sftp.acme.example.com ([203.0.113.10]:22) using "publickey".

Now the failure shapes:

  • No key was offered at all. The -v output jumps straight to Next authentication method: password. The client cannot find the key. The path may be wrong, or it may be the wrong user's home directory. In that case, the job runs as a different account than the one you tested with. Or the private key's permissions are too open and the client refuses it, printing Permissions 0644 for '/home/svc_feed/.ssh/id_ed25519' are too open. Pass the key explicitly with -i to confirm.
  • A key was offered and declined. You see Offering public key followed by Authentications that can continue: publickey and no Server accepts key. The server does not have the matching public key in the account's authorized_keys — or has a different key, or the key was added to the wrong account. Compare fingerprints: ssh-keygen -lf /home/svc_feed/.ssh/id_ed25519.pub on the client, and the same command against the server's authorized_keys file. If they differ, someone is looking at the wrong key.
  • The server refuses to read the key file. The client sees a declined key; the server log says why: Authentication refused: bad ownership or modes for directory /home/feed_acme. The SSH daemon insists that the home directory, .ssh, and authorized_keys are owned by the user and not writable by anyone else. A recursive chmod by a helpful colleague is the usual cause; chmod 700 ~/.ssh && chmod 600 ~/.ssh/authorized_keys the usual fix.
  • The key type is no longer accepted. Unable to negotiate ... no matching host key type found or a server log entry about a disallowed key algorithm: an older key type has been disabled on one side. Generate a modern key and distribute it.
  • The passphrase prompt nobody can answer. The key is fine but protected by a passphrase, and the scheduled job has no agent or terminal to supply it. Interactively it works; unattended it fails. Use an agent, or an unencrypted key stored with tight permissions and a documented owner.

Windows adds one wrinkle. The OpenSSH server on Windows reads administrators' keys from a shared administrators_authorized_keys file rather than the user's profile. The shared file has equally strict permissions. Placing and rotating keys correctly is the subject of our authorized keys distribution guide.

Cause 5: The host key changed

This one runs in the opposite direction. Every SSH server has a host key that identifies it. Every client remembers the host keys it has seen. OpenSSH stores them in a known_hosts file. Graphical Windows clients store them in the registry or a session setting. If the server presents a different key from the one remembered, the client stops:

@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@
@    WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED!     @
@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@
...
Host key verification failed.

Scripted clients skip the banner and print only the last line. So a job that has run for years suddenly fails with Host key verification failed. The server log shows nothing at all — because the client disconnected before attempting to authenticate. That absence in the server log is the fingerprint. The client hung up first, politely, and the server never learned why.

The usual innocent cause is the partner rebuilding or migrating their server without carrying the host key across. The uncommon, serious cause is that you are being connected to a different machine than you think. So the fix is never "just delete the old entry." Confirm the new fingerprint with the partner through a channel other than the connection itself — their onboarding sheet, a phone call. Then, and only then, replace the stored key. Our article on host keys and known_hosts walks through verification and replacement for each client type. I have seen the innocent cause many times and still make the phone call, because the call is cheap and the alternative is not.

Cause 6: A partner-side allowlist

Many partners accept connections only from source addresses they have pre-approved. When your address changes, the partner's server sees a stranger. The cause may be a new internet connection, a firewall replacement, or a job moved to a different server. It may be a VPN that now routes the partner's traffic differently. Depending on how the allowlist is enforced, you see one of three things:

  • A connection that is accepted and then closed before any banner: kex_exchange_identification: Connection closed by remote host, or Connection closed by 203.0.113.10 port 22. This is enforcement at the SSH daemon.
  • A normal login is rejected even though the credential is correct. Sometimes you see an explicit 530 Access denied or, in the partner's log, User feed_acme not allowed because not listed in AllowUsers — enforcement by account-plus-address rules.
  • A silent timeout, if the allowlist lives on the partner's firewall — which looks like layer one, and is why the previous article told you to compare two locations.

The test: find the public address your job's traffic actually leaves from. Your network team knows, and the partner's log shows it. On a machine with several interfaces, Test-NetConnection's SourceAddress shows which one is used. Compare the public address with the address on the partner's onboarding sheet. If they differ, the fix is an email to the partner, not a password reset. The partner's server is not being difficult; it has never met you. Keeping those sheets current is part of partner credentials and security. The wider habit of telling partners about address changes before they bite is covered in firewalls and partner coordination.

Remember: the server log tells you what the server thought happened. Invalid user means the name is wrong. Failed password means the name is right. bad ownership or modes means the key was never read. No entry at all means the client gave up first — a host-key or allowlist problem. Read that line before you touch anything.

The Decision Table

Put the client's output next to the server's log line and the cause usually names itself:

Client shows Server log shows Cause One change to make
Permission denied, or 530 after PASS Invalid user Wrong username Correct the name in the job; nothing else
Permission denied, or 530 after PASS Failed password, one per run Wrong password (stale, mangled, or mistyped) Re-enter the vault's password in the job and test manually
Correct password still fails Burst of failures, then a lock or block entry Locked account or blocked address Fix the source of the failures, then unlock
Correct password fails from a specific date expired password or account-expired entry Expiry Extend or rotate per policy; move the job to a key
Permission denied (publickey), key offered Nothing, or bad ownership or modes Key not installed, or unreadable Compare fingerprints; fix ownership and modes
Host key verification failed Nothing at all Host key changed Verify the new fingerprint out of band, then replace the stored key
Closed before banner, or correct login refused Refused by address rule, or nothing Partner allowlist Send the partner your current source address

Reading the Server Side When You Own the Server

If the failing login is against your server, you have the better half of the evidence. On Linux, the SSH daemon writes to the system journal or the authentication log. journalctl -u ssh --since "02:00" --until "02:30" or grep feed_acme /var/log/auth.log gets you the lines above. An FTP daemon writes its own log with the same facts in its own words. On Windows, the OpenSSH server logs to the event log under its own channel. A dedicated transfer server keeps an activity log. Sysax Multi Server, for example, records every login attempt with the account name, the client's address and the outcome. So a burst of failures from one address or a single unknown account is visible at a glance. Whichever log you read, filter by the source address from your fact sheet first, then by time. A busy server records hundreds of unrelated attempts an hour, most from strangers. Recognizing those strangers is the subject of monitoring authentication attacks.

A Worked Example

The nightly upload to ACME fails at 02:10 with feed_acme@sftp.acme.example.com: Permission denied (publickey). It worked yesterday. The first instinct — regenerate the key and send the partner a new public key — would take a day of partner turnaround. Instead: reproduce as the service account with -v. The output shows Offering public key: /home/svc_feed/.ssh/id_ed25519, then Authentications that can continue: publickey. A key was offered and declined; the client side is fine. You ask the partner for their log at 02:10 and receive one line: Authentication refused: bad ownership or modes for directory /home/feed_acme. Their weekend storage migration reset directory permissions. One chmod on their side, the job runs at 09:00, and nobody regenerated anything. The server had known why all along. Someone finally asked it.

Moving Up the Stack

Once the login succeeds, this layer is done. But note in the write-up which of the six causes it was, because authentication failures recur. Expired passwords expire again, and rotated keys need rotating again. Expiry is a feature; it will feel like a bug on schedule. The prevention is process, laid out in service account hygiene.

The next article, Layer Three: Permission Denied, Decoded, takes the case where the login is accepted and one specific operation is refused. If your problem turned out to be that the client never reached the server at all, step back to Layer One. And for the broader picture of how the authentication methods themselves compare, see authentication methods compared.

Frequently Asked Questions

What does the list in brackets after "Permission denied" mean?
It is the set of authentication methods the server is still willing to accept. "(publickey)" alone means passwords are disabled for this account, so a password-based job can never work there. "(publickey,password)" means both are allowed and whichever you tried was rejected.
The password works when I type it but fails in the script. Why?
Almost always quoting or whitespace: a special character interpreted by the shell, a trailing newline in a password file, or a different character encoding. Print the length of the password the script actually reads and compare it with the vault's. Better still, move the job to key-based authentication.
Should I just delete the known_hosts entry when the host key changes?
Not until you have confirmed the new key's fingerprint with the server's owner through a separate channel. A changed host key is usually a rebuild, but it is also exactly what an interception attack looks like. Verify first, then replace the entry.
How can I tell a lockout from a wrong password?
A wrong password produces a single failure per run in the server log. A lockout shows a burst of failures followed by a lock or block entry, after which a known-correct password fails too. If unlocking works for a few minutes and then the account locks again, a job somewhere is still using the old credential.
The partner says our login attempts never reach them. What now?
Then it is not an authentication problem. Either the client is refusing to proceed or the traffic is dropped before the server sees it. If the client refuses, check for a host key change message. If the traffic is dropped, the cause is an allowlist on their firewall, or a path problem. Confirm your current public source address against what the partner has on file.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.