See the Apache Tika documentation (Using Tika > Tika Server) for full usage.
$ java -jar tika-server/tika-server-standard/target/tika-server-standard-<version>.jar --help
usage: tikaserver
-?,--help this help message
-c,--config <arg> tika-config file
-h,--host <arg> host name (default = localhost, use * for all)
-i,--id <arg> id for this server, written to the startup log
-p,--port <arg> listen port (default = 9998)
Everything beyond host, port and id is configured in the tika-config JSON file passed with -c, not on the command line.
Assuming you have Docker installed, you can use a prebuilt image:
docker run -d -p 127.0.0.1:9998:9998 apache/tika
This will load Apache Tika Server and expose its interface on:
http://localhost:9998
Note the 127.0.0.1: prefix. Unlike the jar, which binds localhost by default, the Docker images start the server with -h 0.0.0.0, so publishing the port without an explicit interface exposes it on every interface of the host. tika-server performs no authentication and parses untrusted files; only expose it on a trusted, access-controlled network. See the Tika Security Model.
You may also be interested in the https://github.com/apache/tika-docker project which provides prebuilt Docker images.
To run as a service on Linux you need to run the install_tika_service.sh script.
Assuming you have the binary distribution tika-server-standard-<version>.zip, you can extract the install script via:
unzip -j tika-server-standard-<version>.zip bin/install_tika_service.sh
and then run the installation process (as root) via:
./install_tika_service.sh ./tika-server-standard-<version>.zip
Usage examples from command line with curl utility:
Extract Markdown (the default output of bare /tika):curl -T price.xls http://localhost:9998/tika
Extract plain text:curl -T price.xls http://localhost:9998/tika/text
Extract Markdown with mime-type hint:curl -v -H "Content-type: application/vnd.openxmlformats-officedocument.wordprocessingml.document" -T document.docx http://localhost:9998/tika
Get all document attachments as ZIP-file:curl -v -T Doc1_ole.doc http://localhost:9998/unpack > /var/tmp/x.zip
Extract metadata as JSON (the default):curl -T price.xls http://localhost:9998/meta
Extract metadata as CSV:curl -T price.xls -H "Accept: text/csv" http://localhost:9998/meta
Detect media type from CSV format using file extension hint:curl -X PUT -H "Content-Disposition: attachment; filename=foo.csv" --upload-file foo.csv http://localhost:9998/detect
200 - Ok204 - No content (for example when we are unpacking file without attachments)400 - Bad request (unknown or invalid handler type, or a reserved/unknown fetcher or emitter was named)403 - Forbidden (per-request configuration was supplied but allowPerRequestConfig is off)413 - Payload too large (the request body exceeds maxRequestSizeBytes, or a pipes payload limit was exceeded)422 - Unparsable document of known type (password protected documents and unsupported versions like Biff5 Excel)429 - Too many requests (all forked workers were busy for longer than maxWaitForClientMillis; retry with backoff)500 - Internal error503 - Service unavailable (the forked worker hit a timeout, ran out of memory, or crashed)