
TensorFlow Serving is a flexible, high-performance serving system for machine learning models designed for production environments. This guide shows how to train a simple model and serve this model with TensorFlow Serving using Docker. It also shows how to serve many models simultaneously.
Use Vultr One-Click Docker to deploy your Docker instance. This guide uses the Ubuntu version.
After the deployment, log in on the server via SSH
Update the machine and reboot the machine to apply the updates
It is recommended to use virtual environments while working on Python projects. Install the package to use Python venv module or use any other method of your preference.
On Ubuntu 20.04:
On Ubuntu 22.04:
After rebooting, switch to docker user and download the TensorFlow Serving image.
Train a model using the Fashion-MNIST dataset which consists of 60k training examples and 10k test examples with ten different classes, where each sample is a 28x28 pixels grayscale image of a piece of clothing.
This guide does not detail the model training or do any parameter tuning. The focus is on TensorFlow Serving rather than the model training.
Set up a virtual environment and install the necessary Python packages.
Create a file model.py with the following code, which loads and prepares the data.
Build a CNN with a single convolutional layer (8 filters of size 3x3) and an output layer with ten units.
Train the model.
Save the model.
Run the code in model.py.
You can visualize the model information with the command saved_model_cli.
Start the server using Docker.
Create a file query.py with the following code, which loads the data, sets the query header, and defines query data as the first example in the test dataset.
Make the query and get the response. Remember to change the address to your server IP or FQDN.
The predictions variable contains a list with the probabilities of the object query belonging to each of the 10 classes. The greatest of these values is the classification made by the model. The function np.argmax returns the model classification from the predictions variable.
Run query.py, and you should see the result of this query.
You can also query for more than one example at each request. Change the data variable to, for example, the five first samples in test data and the output of the query.py to show the classification of each of the five queries.
Running query.py again gives the result of the query with five examples.
TensorFlow Serving is also capable of serving more than one model simultaneously. This time, use some of the pre-trained models available in the TensorFlow Serving GitHub repository.
Install Git if it is not installed yet.
And clone the repository.
Create a file named models.config with the content below, which configures two models defining the name, directory relative to the Docker container running your server, and model platform for each one.
The models compute the half part of a number and add two and three, respectively.
Start the server loading this configuration file and mounting the volume with the pre-trained models.
You can view each model status information with an HTTP GET request on the browser or using curl.
And request predictions for both models with an HTTP POST request using curl by running:
This guide covered how to train a TensorFlow model and deploy TensorFlow Serving to serve this and many models simultaneously with Docker.
To learn more about TensorFlow Serving, go to the developer documentation pages:
0 Comments
Be the first to comment and share your perspective with the community.