Using EvoPdf Next for .NET in AWS Lambda

EvoPdf Next for .NET is a library that can be integrated into AWS Lambda functions to create and process PDF documents.

You can create PDF documents, convert HTML, Word, Excel, RTF and Markdown documents to PDF, extract text and images from existing PDF documents, perform text search operations on PDF documents and convert PDF pages to images.

Platform Compatibility

EvoPdf Next for .NET runs in Lambda functions deployed as container images, built on the Lambda base image for .NET 10, which is based on Amazon Linux 2023, on the x86_64 architecture and on the arm64 architecture with AWS Graviton processors.

A Lambda function deployed as a zip archive is limited to 250 MB of unpacked code, which is less than the size of the HTML to PDF Converter runtime, so the function is packaged as a container image. The image contains the .NET runtime, the function, the library runtime, the system packages and the fonts and needs no setup when the function runs. Lambda runs the function with a non-root user and allows writes only in the /tmp folder; the image built as described below takes care of both.

The library targets .NET Standard 2.0, making it usable in any .NET Core application that supports this standard.

Create the Lambda Function Project

Create a .NET 10 class library project for the function and add references to the Lambda packages and to the EvoPdf.Next.HtmlToPdf.Linux NuGet package. To use other components of the library, reference the EvoPdf.Next.Linux metapackage instead, which installs all the components. You can find more details about the available packages in the Getting Started on Linux documentation section.

Project file
<Project Sdk="Microsoft.NET.Sdk">
  <PropertyGroup>
    <TargetFramework>net10.0</TargetFramework>
    <ImplicitUsings>enable</ImplicitUsings>
    <GenerateRuntimeConfigurationFiles>true</GenerateRuntimeConfigurationFiles>
    <AWSProjectType>Lambda</AWSProjectType>
  </PropertyGroup>
  <ItemGroup>
    <PackageReference Include="Amazon.Lambda.Core" Version="2.5.0" />
    <PackageReference Include="Amazon.Lambda.APIGatewayEvents" Version="2.7.1" />
    <PackageReference Include="Amazon.Lambda.Serialization.SystemTextJson" Version="2.4.4" />
    <PackageReference Include="EvoPdf.Next.HtmlToPdf.Linux" Version="14.86.0" />
  </ItemGroup>
</Project>

The function below answers requests received through a function URL or an API Gateway HTTP API. It converts the web page given in the url query string parameter (or an HTML string when the parameter is missing) and returns the PDF document. Lambda transfers binary responses encoded in base64.

Function.cs
using Amazon.Lambda.APIGatewayEvents;
using Amazon.Lambda.Core;
using EvoPdf.Next;

[assembly: LambdaSerializer(typeof(Amazon.Lambda.Serialization.SystemTextJson.DefaultLambdaJsonSerializer))]

namespace PdfFunction;

public class Function
{
    public APIGatewayHttpApiV2ProxyResponse FunctionHandler(APIGatewayHttpApiV2ProxyRequest request, ILambdaContext context)
    {
        string url = null;
        request.QueryStringParameters?.TryGetValue("url", out url);

        HtmlToPdfConverter converter = new HtmlToPdfConverter();

        byte[] pdf = string.IsNullOrEmpty(url)
            ? converter.ConvertHtml("<b>Hello World</b> from EvoPdf Next!", null)
            : converter.ConvertUrl(url);

        return new APIGatewayHttpApiV2ProxyResponse
        {
            StatusCode = 200,
            Body = Convert.ToBase64String(pdf),
            IsBase64Encoded = true,
            Headers = new Dictionary<string, string> { ["Content-Type"] = "application/pdf" }
        };
    }
}

Build the Container Image

Publish the function for Linux. The publish folder contains the function, its dependencies and the library runtime:

 
dotnet publish -c Release -r linux-x64 --self-contained false -o publish

Create the Dockerfile below in the project folder. It starts from the Lambda base image for .NET 10, installs the system packages required by the HTML to PDF Converter and the DejaVu font families, copies the publish folder and gives execute permission to the runtime files, which a publish folder created on Windows does not keep. The DejaVu fonts render the text of pages that do not use web fonts, with DejaVu Sans set as the sans-serif font; you can add other font packages or copy your own font files into the image. The HOME variable points the working files of the converter to /tmp, the writable folder of the function.

Dockerfile
FROM public.ecr.aws/lambda/dotnet:10

# Install the packages required by the HTML to PDF Converter and the DejaVu font families
RUN dnf install -y nss at-spi2-atk cairo pango dejavu-sans-fonts dejavu-serif-fonts dejavu-sans-mono-fonts && dnf clean all

# Set DejaVu Sans as the sans-serif font
RUN printf '<fontconfig><alias><family>sans-serif</family><prefer><family>DejaVu Sans</family></prefer></alias></fontconfig>' > /etc/fonts/local.conf

# Copy the published function
COPY publish/ ${LAMBDA_TASK_ROOT}

# Ensure execute permissions for the HTML to PDF Converter runtime
RUN chmod +x ${LAMBDA_TASK_ROOT}/evopdf_runtimes/linux-x64/native/evopdf_loadhtml

# The only writable folder of a Lambda function
ENV HOME=/tmp

CMD ["PdfFunction::PdfFunction.Function::FunctionHandler"]

The CMD line names the handler as assembly, type and method. If the function also uses the PDF Processor component, installed by the EvoPdf.Next.Linux metapackage, add the same permission line for the evopdf_runtimes/linux-x64/native/evopdf_pdfprocessor file.

For a function that targets .NET 8, set net8.0 as the target framework and start from the public.ecr.aws/lambda/dotnet:8 image, which is also based on Amazon Linux 2023; the packages are the same.

Build the image for the x86_64 architecture. Lambda accepts images with a single manifest, so the attestations that recent Docker versions add by default are disabled:

 
docker build --platform linux/amd64 --provenance=false --sbom=false -t pdf-function .

For an arm64 function, reference the EvoPdf.Next.HtmlToPdf.Linux.Arm64 package in place of EvoPdf.Next.HtmlToPdf.Linux, publish for the linux-arm64 runtime, replace linux-x64 with linux-arm64 in the path of the Dockerfile and build the image for the arm64 architecture:

 
dotnet publish -c Release -r linux-arm64 --self-contained false -o publish
docker build --platform linux/arm64 --provenance=false --sbom=false -t pdf-function .

On a computer with an x64 processor Docker builds the arm64 image by emulation, which makes the installation of the packages slower.

Push the Image and Create the Function

Lambda loads the image from a private repository in Amazon Elastic Container Registry in the same region as the function. With the AWS CLI configured for your account, create the repository, sign in to the registry and push the image; replace the account number and the region with your own:

 
aws ecr create-repository --repository-name pdf-function --region eu-central-1
aws ecr get-login-password --region eu-central-1 | docker login --username AWS --password-stdin 123456789012.dkr.ecr.eu-central-1.amazonaws.com
docker tag pdf-function:latest 123456789012.dkr.ecr.eu-central-1.amazonaws.com/pdf-function:latest
docker push 123456789012.dkr.ecr.eu-central-1.amazonaws.com/pdf-function:latest

In the Lambda console choose Create function and Container image, select the image from the repository and, under Additional settings, the architecture of the image: x86_64, the default, or arm64. The function needs the basic execution role, which the console creates, with the permission to write its logs. After the function is created, set its configuration:

  • Memory: HTML to PDF conversion can be resource-intensive, depending on the complexity of the content; Lambda allocates processor power in proportion to the memory. Configure at least 2048 MB.

  • Timeout: the first conversion in a new execution environment also starts the converter and takes longer than the following ones. Configure a timeout of at least 2 minutes.

  • Function URL: to call the function over HTTPS, create a function URL in the Configuration tab. The AWS_IAM authentication type accepts signed requests only; NONE makes the URL public.

The same settings are available from the AWS CLI with the aws lambda create-function, aws lambda update-function-configuration and aws lambda create-function-url-config commands.

Run the Function

Open the function URL in a browser to receive the PDF document, or add ?url=https://www.example.com to convert a web page. The first request in a new execution environment takes longer, while Lambda starts the environment and the function starts the converter; the following requests served by the same environment convert at the usual speed. Lambda keeps an environment for a while after the last request and starts a new one when the requests resume.

As a reference, in the Europe (Frankfurt) region, with an x86_64 function of 3008 MB and an arm64 function of 2048 MB, the first request of a new environment took 7 to 10 seconds. In the same environment a one-page HTML string then converted in 0.2 to 0.3 seconds and a six-page web site home page in 2 to 3.5 seconds, with less than 650 MB of memory used.

To deploy a new version of the function, push the new image and choose Deploy new image in the Image section of the function page.

The response of a Lambda function is limited to 6 MB. The PDF document travels encoded in base64, so documents up to about 4.5 MB fit in the response. For larger documents, save the PDF to an Amazon S3 bucket and return its link.

Troubleshooting

The output and the exceptions of the function are written to the CloudWatch log group of the function, available from the Monitor tab. The REPORT line of every invocation shows its duration and the memory used.

If the function cannot be created from the image, with an error about the manifest or the media type of the image, build the image again with the --provenance=false --sbom=false options.

If the function URL returns Internal Server Error and the log shows that the invocation ended before the conversion, check the memory and the timeout of the function: a new function starts with 128 MB of memory and a timeout of 3 seconds.

See Also