Even the best automated solutions rely on recognising someone who is speaking clearly and precisely, ensuring that every word is well spaced and completely clear.
There are no automated solutions that can universally do a good job of transcribing natural speech from people who aren't specifically "speaking to be recognised", if that makes sense.
Maybe one day, but not yet. It's a problem waiting to be solved, so the reward for the first who can really crack it will be substantial.
What software/sdk's have you used?