Also known as the build stage of the SDLC, coding focuses on the writing and programming of a system. The Zones in this category take a hands-on approach to equip developers with the knowledge about frameworks, tools, and languages that they can tailor to their own build needs.
A framework is a collection of code that is leveraged in the development process by providing ready-made components. Through the use of frameworks, architectural patterns and structures are created, which help speed up the development process. This Zone contains helpful resources for developers to learn about and further explore popular frameworks such as the Spring framework, Drupal, Angular, Eclipse, and more.
Java is an object-oriented programming language that allows engineers to produce software for multiple platforms. Our resources in this Zone are designed to help engineers with Java program development, Java SDKs, compilers, interpreters, documentation generators, and other tools used to produce a complete application.
JavaScript (JS) is an object-oriented programming language that allows engineers to produce and implement complex features within web browsers. JavaScript is popular because of its versatility and is preferred as the primary choice unless a specific function is needed. In this Zone, we provide resources that cover popular JS frameworks, server applications, supported data types, and other useful topics for a front-end engineer.
Programming languages allow us to communicate with computers, and they operate like sets of instructions. There are numerous types of languages, including procedural, functional, object-oriented, and more. Whether you’re looking to learn a new language or trying to find some tips or tricks, the resources in the Languages Zone will give you all the information you need and more.
Development and programming tools are used to build frameworks, and they can be used for creating, debugging, and maintaining programs — and much more. The resources in this Zone cover topics such as compilers, database management systems, code editors, and other software tools and can help ensure engineers are writing clean code.
Six Degrees of Ayrton Senna: Learn Neo4j by Connecting 75 Years of Formula 1
The Silent Container Death: A TCP Dial That Never Times Out
A self-healing SQL pipeline should not mean autonomous SQL generation followed by privileged execution. In production, the safer interpretation is narrower, where a language model proposes a repair, while deterministic controls decide whether that repair is syntactically valid, semantically plausible, operationally safe, and eligible for execution. This distinction matters because the same mechanism that corrects a renamed column can also generate an unintended DELETE, widen a join, or scan an unexpectedly large dataset. Structured-output features can constrain an LLM response to a defined schema, but schema conformance is not equivalent to database correctness or authorization. OpenAI’s Structured Outputs is designed to make generated output conform to supplied JSON Schemas, and it does not validate SQL semantics or execution safety. Treat the Failure as Evidence, Not Merely a Prompt The repair loop should begin by classifying the failure before any model call. Parser errors, missing relations, unknown columns, type mismatches, permission failures, timeouts, cardinality explosions, and upstream freshness problems require different responses. A permission error should not trigger SQL rewriting, while an unknown-column error may justify metadata inspection. This gate keeps deterministic failure classes deterministic. Consider a pipeline that previously executed: SQL SELECT customer_id, customer_segment FROM analytics.customer_profile WHERE active = TRUE; An upstream migration renames customer_segment to segment_name. The database returns an unknown-column error. The repair service should collect the failed statement, SQL dialect, error code, schema version, referenced objects, and recent catalog changes. Metadata inspection can then establish that customer_profile still exists, customer_segment no longer exists, and segment_name appeared in the latest schema version. That evidence is stronger than asking a model to infer a replacement from exception text alone. Schema drift detection should therefore precede generation. Catalog snapshots can be hashed and compared between successful and failed runs. Candidate mappings can incorporate data type compatibility, nullability, lineage metadata, column comments, and migration records. The LLM then receives only evidence relevant to the suspected failure class. Candidate generation should return a structured proposal rather than free-form SQL. A production contract can require the proposed SQL, repair category, changed identifiers, evidence references, and assumptions: Python def generate_candidate(failure, metadata): return llm.generate( schema=RepairProposal, context={"failure": failure, "metadata": metadata}, constraints={"max_statements": 1, "allow_dml": False} ) The allow_dml flag is an application policy, not an instruction trusted merely because it appears in a prompt. Structured generation narrows output shape, while authorization remains outside the model. Anthropic’s evaluation guidance similarly emphasizes measurable success criteria and testable thresholds rather than treating model behavior as inherently reliable. Parse the Candidate Before the Database Sees It String matching is too weak for SQL safety. A rule such as "DELETE" not in sql.upper() can miss nested statements and dialect-specific constructs. The candidate should be parsed into an abstract syntax tree using the correct database dialect. SQLGlot parses SQL into expression trees and supports multiple dialects, enabling structural inspection before execution. A validator can reject statement types outside an allowlist and verify referenced tables and columns against current metadata: Python def validate_ast(sql, dialect, catalog): tree = parse_one(sql, dialect=dialect) if tree.find(Delete) or tree.find(Update) or tree.find(Insert): raise PolicyViolation("mutating statement rejected") for table in tree.find_all(Table): catalog.require_table(table.name) for column in tree.find_all(Column): catalog.require_column(column.table, column.name) return tree AST validation should also enforce tenant boundaries, prohibited schemas, mandatory predicates, join limits, and function restrictions. A generated query can be syntactically valid yet unsafe because an omitted filter changes a targeted lookup into a full table operation. Semantic checks therefore need context from the original successful query, expected output columns, and data quality assertions. The repaired query may be: SQL SELECT customer_id, segment_name AS customer_segment FROM analytics.customer_profile WHERE active = TRUE; Preserving the original output alias matters because downstream consumers may depend on customer_segment even though the physical source column changed. A repair that only substitutes the new identifier could restore execution while silently breaking the pipeline contract. Dry Runs Should Prove More Than Syntax A candidate that survives static validation still should not immediately reach production data. Database-native validation can catch failures that an AST cannot. BigQuery dry runs validate query structure and estimate bytes processed without executing the query, making them useful for rejecting unexpectedly expensive repairs. Google also documents that successful dry runs do not guarantee successful runtime execution and that multi-statement dry runs have special limitations. A guarded validation step can combine dry run results with policy limits: Python def dry_run(candidate): result = warehouse.validate(candidate.sql) if result.bytes_scanned > MAX_BYTES: raise PolicyViolation("scan budget exceeded") if result.output_schema != candidate.expected_schema: raise ContractViolation("output schema changed") return result For engines without native dry run support, a read-only transaction, isolated replica, sandbox database, or planner-only operation can provide a safer boundary. PostgreSQL supports read-only transaction modes that prevent changes to non-temporary tables, adding a database-enforced control rather than relying only on application logic. Operational safety also depends on continuous monitoring after a repair is deployed. Query latency, row counts, null rates, schema changes, and downstream data quality indicators can reveal subtle regressions that static validation may miss. Post execution monitoring therefore provides another deterministic checkpoint, allowing suspicious behavior to trigger rollback or human review immediately. Confidence should come from independently observable signals, not from an LLM declaring confidence in its own answer. A score can combine schema evidence, AST-policy results, dry run success, output-schema stability, and historical repair success: Python def repair_score(signals): return ( 0.30 * signals.schema_evidence + 0.25 * signals.ast_validation + 0.20 * signals.dry_run + 0.15 * signals.contract_match + 0.10 * signals.history ) Weights should be calibrated against labeled historical failures. LLM-based judging can contribute a secondary semantic signal, but it should not authorize execution. G-Eval shows that model-based evaluators can correlate with human judgments while also identifying evaluator bias as a concern. A useful inference from RAGAS is that component-level measurements are more diagnosable than one opaque score; the same principle fits SQL repair by keeping schema, syntax, execution, and contract evidence separately observable. Bound Autonomy and Make Every Repair Reversible Self-healing becomes dangerous when retries are unbounded. A failed candidate should feed only new deterministic evidence into the next attempt, such as a parser error or dry-run diagnostic, and the loop should stop after a small configured limit. Repeated failures, low confidence, ambiguous schema mappings, contract changes, or any proposed mutation should escalate to human review. LangSmith distinguishes offline evaluation from online production evaluation and describes a feedback loop in which production failures become future evaluation cases and the same operational pattern fits SQL repair systems. Every attempt should produce an immutable audit record containing the original SQL hash, failure evidence, metadata version, model and prompt version, candidate hash, validation outcomes, confidence signals, execution identity, and final disposition. That record supports debugging and regression testing of future repair policies. Production monitoring should also track changes in failure distributions and repair success rates, as Google’s MLOps guidance treats monitoring as a trigger for new experimentation when production quality degrades. Automatic mutation deserves a higher bar than automatic read-only repair. When writes are permitted, idempotency keys should prevent duplicate side effects across retries, and execution should remain transactional whenever supported. AWS reliability guidance recommends idempotency for database insert, update, and delete operations. Transaction savepoints and rollback mechanisms add another containment layer, as PostgreSQL savepoints allow effects after a savepoint to be selectively discarded without abandoning the entire transaction. A production-grade self-healing SQL pipeline is not an autonomous database administrator implemented with a prompt. It is a controlled repair system in which probabilistic generation is surrounded by deterministic evidence collection, AST inspection, schema validation, database-enforced dry runs, calibrated confidence thresholds, bounded retries, auditability, and explicit escalation. The safest design grants the LLM authority to propose change, not authority to approve or execute it. With that separation preserved, LLMs can reduce recovery time for routine SQL failures while the database, policy engine, and human review path retain control over production state.
Most "RAG over PDFs" pipelines have a step nobody talks about much: something has to turn a scanned invoice, a multi-column contract, or a photographed receipt into text a model can actually reason over. On Microsoft's stack, that something is usually the Document Intelligence SDK, formerly Form Recognizer, and it's worth understanding on its own terms rather than treating it as a black box that happens before the interesting part starts. This is a hands-on deep dive into that SDK specifically. Not a tour of every Foundry Tools SDK — Vision and Speech and Content Safety each deserve their own treatment, but a real build using Document Intelligence: extracting layout as clean markdown, pulling structured fields out of a known document type, classifying documents before routing them, and training a custom extraction model on your own labeled data. The Mental Model First Two clients, and three kinds of model, cover almost everything this SDK does: DocumentIntelligenceClient runs analysis. Every call goes through one method, begin_analyze_document, and a model_id parameter decides what kind of analysis happens. It's a long-running operation, so every call returns a poller.DocumentIntelligenceAdministrationClient manages models. This is where you build custom extraction models and classifiers, list what's already been trained, and delete what you don't need anymore.Prebuilt models (prebuilt-layout, prebuilt-invoice, prebuilt-receipt, prebuilt-idDocument, prebuilt-read, and others) handle common, well-known document shapes out of the box. No training required.Custom extraction models, trained on your own labeled documents, handle document types nobody prebuilt a model for: your specific contract template, your specific intake form.Classifiers solve a different problem entirely: given a document of unknown type, which model should even look at it? This matters more than it sounds like it should, since most real document pipelines receive a mix of types, not one known shape. Prerequisites A Document Intelligence resource (or a multi-service Foundry resource, which includes it), giving you an endpoint and either an API key or Entra ID access.Python 3.9+ with the SDK installed. Python pip install azure-ai-documentintelligence azure-identity Python from azure.ai.documentintelligence import DocumentIntelligenceClient from azure.core.credentials import AzureKeyCredential endpoint = "https://YOUR-RESOURCE.cognitiveservices.azure.com" client = DocumentIntelligenceClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) For anything past local experimentation, swap the key for DefaultAzureCredential and an RBAC role scoped to the resource, the same pattern every other Foundry-adjacent SDK in this series has used. Step 1: Layout Extraction, Straight to Markdown This is the single most useful call in the whole SDK if your end goal is feeding documents into a RAG pipeline. prebuilt-layout doesn't just extract text; it understands headings, tables, and section structure, and it can hand all of that back as GitHub-flavored markdown instead of a flat text blob. Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentContentFormat with open("contract.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), output_content_format=DocumentContentFormat.MARKDOWN, ) result = poller.result() print(result.content[:500]) result.content is now a markdown string, headings as #, tables as GFM pipe tables, page structure preserved. That matters more than it sounds like it should: a table flattened into plain text loses its row and column relationships, and a model reasoning over that text has to reconstruct structure it was never actually given. Markdown output keeps the structure intact. Step 2: Pulling Structured Fields From a Known Document Type For document types Document Intelligence already knows, invoices are the clearest example; you get named fields back with a confidence score per field, not just raw text. Python with open("invoice.pdf", "rb") as f: poller = client.begin_analyze_document("prebuilt-invoice", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: vendor = doc.fields.get("VendorName") total = doc.fields.get("InvoiceTotal") if vendor: print(f"Vendor: {vendor.value_string} (confidence: {vendor.confidence:.2f})") if total: print(f"Total: {total.value_currency.amount} (confidence: {total.confidence:.2f})") That confidence score isn't decoration. It's the field you should actually branch on in production code; more on that in the production section below. Step 3: Add-On Capabilities You'll Want More Often Than the Docs Suggest A few optional capabilities aren't on by default, since they add processing cost, but are worth turning on deliberately rather than discovering you needed them after the fact: Python from azure.ai.documentintelligence.models import AnalyzeDocumentRequest, DocumentAnalysisFeature with open("shipping-label.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", AnalyzeDocumentRequest(bytes_source=f.read()), features=[DocumentAnalysisFeature.BARCODES, DocumentAnalysisFeature.FORMULAS], ) BARCODES extracts barcode and QR code payloads directly, useful for shipping labels and inventory documents where the barcode carries the actual identifier the text doesn't repeat. FORMULAS pulls out mathematical expressions as LaTeX, relevant if you're processing scientific or financial documents where a formula matters more than the surrounding prose. There's also a high-resolution mode for documents where small print matters, at the cost of slower processing. Step 4: Build a Classifier to Route Mixed Document Types Real intake pipelines rarely receive one document type. A classifier solves the "what am I even looking at" problem before you commit to an extraction model. Python from azure.ai.documentintelligence import DocumentIntelligenceAdministrationClient from azure.ai.documentintelligence.models import ( BuildDocumentClassifierRequest, ClassifierDocumentTypeDetails, AzureBlobContentSource, ) admin_client = DocumentIntelligenceAdministrationClient(endpoint=endpoint, credential=AzureKeyCredential("YOUR-KEY")) poller = admin_client.begin_build_classifier( BuildDocumentClassifierRequest( classifier_id="support-doc-classifier", doc_types={ "invoice": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-invoices-container>") ), "contract": ClassifierDocumentTypeDetails( azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-contracts-container>") ), }, ) ) classifier = poller.result() You need at least five sample documents per category to train a classifier at all, and more than that for anything you'd trust in production. Once it's built, classifying an incoming document is a single call: Python with open("unknown.pdf", "rb") as f: poller = client.begin_classify_document("support-doc-classifier", AnalyzeDocumentRequest(bytes_source=f.read())) result = poller.result() for doc in result.documents: print(f"Classified as: {doc.doc_type} (confidence: {doc.confidence:.2f})") Step 5: Build a Custom Extraction Model for Your Own Document Type When a document type isn't invoices, receipts, or any of the other prebuilt shapes, train your own. This needs a set of labeled training documents in Blob Storage, produced through the labeling tool in Foundry's document intelligence studio or programmatically. Python from azure.ai.documentintelligence.models import ( BuildDocumentModelRequest, AzureBlobContentSource, DocumentBuildMode, ) poller = admin_client.begin_build_document_model( BuildDocumentModelRequest( model_id="acme-service-agreement-v1", build_mode=DocumentBuildMode.TEMPLATE, azure_blob_source=AzureBlobContentSource(container_url="<SAS-url-to-training-container>"), description="Extraction model for Acme's standard service agreement template.", ) ) model = poller.result() Two build modes matter here, and they're not interchangeable. TEMPLATE mode is faster to train and works well when your documents follow a consistent visual layout, the same form filled out differently each time. NEURAL mode handles structural variation better, different layouts that still represent the same document type, at the cost of needing more training examples and longer build time. Start with TEMPLATE unless your documents genuinely vary in structure, not just content. One naming constraint worth knowing before you hit it: a custom model ID can't start with prebuilt-, since that prefix is reserved for Microsoft's own models across every resource. Where This Fits in the Bigger Picture This is the detail that trips people up once they've also worked with the Foundry SDK or Agent Framework elsewhere in this series: Document Intelligence doesn't go through your Foundry project endpoint at all. It has its own resource, its own endpoint (resource.cognitiveservices.azure.com), and its own authentication scope. That's what "Foundry Tools SDK" actually means as a category, prebuilt AI services with tool-specific endpoints, distinct from the Foundry SDK's unified project endpoint that Agent Framework and the Responses API build on. The practical upshot is the pipeline most teams actually want: run prebuilt-layout over incoming documents, get markdown back, and hand that markdown to a Foundry IQ Knowledge Base as a File Knowledge Source. Document Intelligence handles turning the PDF into clean, structured text. Foundry IQ handles chunking, embedding, and retrieval on top of it. Neither service needs to know the other exists; they just happen to compose well because Markdown is a reasonable interchange format for both. Production Considerations Before You Commit Don't trust a field just because it came back. A field with a confidence score of 0.41 should not silently flow into a downstream system as if it were as reliable as one scored 0.98. Set a threshold, route low-confidence extractions to human review, and log the confidence distribution over time so a model quietly degrading on a document template change doesn't go unnoticed.Classifier training minimums are a floor, not a target. Five documents per category is what the service requires to build at all. It is not enough to trust a classifier's accuracy in production. Budget for real evaluation data, held out from training, before routing real documents based on classifier output.TEMPLATE vs NEURAL is a real tradeoff, not a default to leave unexamined. Picking NEURAL by default because it sounds more capable means slower training and a higher training-data bar for a benefit you may not need if your documents are already visually consistent.Preview API versions and regional availability move independently of the SDK version. A given SDK release doesn't guarantee every feature is available in every region. Check current regional availability for newer capabilities (certain add-ons, newer prebuilt models) before designing around them.Markdown output is currently scoped to prebuilt-layout. Don't assume other prebuilt or custom models will hand back the same content format; check per-model support before building a pipeline that assumes Markdown everywhere.Cost scales with pages and capability, not just call count. Add-on features like high-resolution mode and custom model training both carry their own cost beyond the base per-page analysis price. Model this before committing to a design that turns on every add-on by default. Where This Leaves You The Document Intelligence SDK is easy to undersell because the interesting part of most AI applications feels like it's happening somewhere else, in the model, in the retrieval layer, in the agent's reasoning. But the quality ceiling of everything downstream is set right here, at the point where a physical or scanned document either does or doesn't become text a model can actually use well. Layout extraction to markdown, confidence-aware field extraction, classifiers for mixed intake, and custom models for your own document shapes cover the large majority of real document-processing needs, and all four are a few lines of SDK code once you know which one you need. The judgment call was never really about the API. It's about matching the right one of these four tools to what's actually in your inbound documents. References Microsoft. "azure-ai-documentintelligence README." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/README.mdMicrosoft Learn. "Document Intelligence layout model." learn.microsoft.com/en-us/azure/ai-services/document-intelligence/prebuilt/layoutMicrosoft. "Migration guide, azure-ai-documentintelligence." Azure SDK for Python. github.com/Azure/azure-sdk-for-python/blob/main/sdk/documentintelligence/azure-ai-documentintelligence/MIGRATION_GUIDE.mdMicrosoft Learn. "Get started with Microsoft Foundry SDKs and endpoints." learn.microsoft.com/en-us/azure/foundry/how-to/develop/sdk-overviewMicrosoft Learn. "What is Foundry IQ?" learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq
Batch processing remains vital because many business operations aren't suited to interactive requests. Tasks such as recalculating prices, reconciling transactions, migrating records, generating reports, processing invoices, reclassifying customers, or applying rules across millions of records may require considerable time. Handling these as standard requests leads to fragile systems, increased user wait times, frequent timeouts, challenging retries, and possible data inconsistencies. A batch model handles large workloads predictably, incrementally, and with control over progress and recovery. Rather than processing a massive operation as a single loop, batch processing uses jobs, steps, chunks, checkpoints, filtering, and restartability. This approach separates long-running data tasks from the user experience while delivering a structured execution model. In this article, we will focus on Jakarta Batch and examine its sustained relevance for modern enterprise applications. Why Batch Processing Still Matters in Enterprise Systems Modern applications offer various methods for background processing, such as message queues, event-driven architectures, schedulers, reactive pipelines, and distributed stream-processing platforms. While each addresses specific needs, batch processing is most effective when operations have a defined start and end, involve a known or discoverable dataset, and require controlled execution, progress tracking, restartability, or periodic processing. Batch processing remains essential in enterprise systems. Workloads such as financial reconciliation, billing, payroll, reporting, data migration, regulatory processing, catalog updates, and large-scale reclassification are still prevalent. In these scenarios, the priority is to process large volumes of work safely and predictably, rather than responding to individual events quickly. Batch provides a model specifically designed for these requirements. How Jakarta Batch Works Jakarta Batch organizes background processing into jobs and steps. A job defines the overall batch operation, while each step represents a specific stage. In chunk-oriented processing, a step follows a simple pipeline: read, process, write, and repeat until it processes all input. The Jakarta Batch runtime manages this lifecycle so application code can focus on reading, transforming, and persisting data. A job is the top-level unit of execution and represents a complete business operation, such as importing records, recalculating customer classifications, processing invoices, or reconciling transactions. Jobs can accept parameters at startup, allowing the same batch definition to run with different inputs or business rules. A job consists of one or more steps, each representing a distinct phase of the workload. Simple jobs may have a single step, while complex processes can use multiple steps in sequence, such as importing data, validating it, and generating a final report. Within a chunk-oriented step, the ItemReader supplies data to the runtime one item at a time, from sources such as a database or file. The reader only retrieves the next item and does not need to know how it will be processed or persisted.The ItemProcessor receives each item and applies business rules, such as validation, transformation, classification, enrichment, or filtering. It may return a modified item or null if the item should be excluded from writing.The ItemWriter receives processed items and persists or exports them. Unlike the reader and processor, which handle items individually, the writer typically receives a group of items from the current chunk. This enables more efficient database or bulk operations. Jakarta Batch adds features around this pipeline to support enterprise workloads. The runtime manages chunk boundaries, transactions, checkpoints, execution status, failures, and restart behavior. Chunk size determines how much work is grouped before a write and checkpoint, making it a key parameter for balancing throughput, memory usage, database cost, and recovery. The core model is straightforward: Job → Step → Read → Process → Write → Repeat Jakarta Batch keeps the business pipeline simple while the runtime manages the execution mechanics needed for reliable, long-running data processing. The Sample: Customer Segmentation with Jakarta Batch This example demonstrates the Jakarta Batch model using an e-commerce customer segmentation scenario. Customers are assigned to tiers such as Bronze, Silver, Gold, and Platinum based on configurable spending thresholds. When thresholds change, the application reevaluates the customer base and updates only customers whose classification has changed. The full application includes MongoDB integration, a Jakarta Faces UI, a preview workflow, validation, and supporting services. The complete source code is available at https://github.com/soujava/mongodb-jakarta-batch. This section focuses on the classes directly involved in Jakarta Batch execution. Starting the Batch Job The application initiates the batch process through CustomerSegmentationService. Unlike the reader, processor, and writer, this class is not a batch artifact. Instead, it is an application service that retrieves Jakarta Batch’s JobOperator from BatchRuntime to start and monitor job executions. Java @ApplicationScoped public class CustomerSegmentationService { public static final String JOB_NAME = "customer-segmentation"; private volatile CustomerSegmentationPolicy currentPolicy; // initialization and status methods omitted public long start(CustomerSegmentationPolicy policy) { if (isRunning()) { throw new IllegalStateException( "A customer segmentation batch is already running"); } Properties parameters = new Properties(); parameters.setProperty( CustomerSegmentationPolicy.JOB_PARAMETER, policy.toJson()); long executionId = BatchRuntime.getJobOperator() .start(JOB_NAME, parameters); currentPolicy = policy; return executionId; } public boolean isRunning() { // implementation omitted } } The key API here is JobOperator, which Jakarta Batch provides as the interface for starting, stopping, restarting, and inspecting jobs. In this example, the segmentation policy is serialized into the job parameters to ensure each execution gets the correct business rules. Reading the Input The first batch artifact, CustomerItemReader, extends Jakarta Batch’s AbstractItemReader to implement a chunk-oriented reader. Java @Named("customerItemReader") @Dependent public class CustomerItemReader extends AbstractItemReader { @Inject private CustomerRepository customerRepository; private List<Customer> customers = List.of(); private int nextIndex; @Override public void open(Serializable checkpoint) { try (Stream<Customer> customerStream = customerRepository.findAll()) { customers = customerStream .sorted(Comparator.comparing(Customer::getId)) .toList(); } nextIndex = checkpoint instanceof Integer index ? index : 0; } @Override public Customer readItem() { if (nextIndex >= customers.size()) { return null; } return customers.get(nextIndex++); } @Override public Serializable checkpointInfo() { return nextIndex; } } These methods are part of the Jakarta Batch reader lifecycle defined by AbstractItemReader. open() prepares the reader and accepts a previous checkpoint if available. readItem() provides the next item to the runtime; returning null indicates there is no more input. checkpointInfo() reports the reader’s current position for checkpointing. For simplicity, this sample loads customers into memory. For larger workloads, the implementation might use pagination or a MongoDB cursor without changing the Jakarta Batch model. Processing Each Customer The next artifact implements Jakarta Batch’s ItemProcessor interface. Java @Named("customerTierProcessor") @Dependent public class CustomerTierProcessor implements ItemProcessor { @Inject @BatchProperty( name = CustomerSegmentationPolicy.JOB_PARAMETER) private String thresholdsJson; private CustomerSegmentationPolicy policy; @PostConstruct void initialize() { policy = CustomerSegmentationPolicy.fromJson( thresholdsJson); } @Override public Customer processItem(Object item) { if (!(item instanceof Customer customer)) { throw new IllegalArgumentException( "Expected a Customer item"); } CustomerTier calculatedTier = policy.tierFor(customer.getTotalSpent()); if (calculatedTier == customer.getTier()) { return null; } return Customer.builder() .id(customer.getId()) .name(customer.getName()) .totalSpent(customer.getTotalSpent()) .tier(calculatedTier) .build(); } } Here the Jakarta Batch contract is explicit: ItemProcessor defines processItem(). The runtime calls that method for every item produced by the reader. The processor applies the segmentation rule and either returns the transformed customer or null. Returning null has a specific meaning in Jakarta Batch: the item is filtered and does not continue to the writer. The @BatchProperty is also part of the Batch integration. It receives the thresholds property defined for this job execution, allowing the processor to reconstruct the CustomerSegmentationPolicy before processing begins. Writing the Results The final artifact extends AbstractItemWriter, Jakarta Batch’s base implementation for writing a chunk. Java @Named("customerItemWriter") @Dependent public class CustomerItemWriter extends AbstractItemWriter { @Inject private CustomerRepository customerRepository; @Override public void writeItems(List<Object> items) { List<Customer> customers = items.stream() .map(this::toCustomer) .toList(); customerRepository.saveAll(customers); } private Customer toCustomer(Object item) { if (item instanceof Customer customer) { return customer; } throw new IllegalArgumentException( "Expected a Customer item"); } } writeItems() is defined by the Jakarta Batch writer contract inherited from AbstractItemWriter. Unlike the processor, which receives one item at a time, the writer receives a collection of processed items. In this case, the collection contains only customers whose classification changed, as the processor has already filtered the others. At this point, the Java components of the pipeline are as follows: Plain Text CustomerItemReader extends AbstractItemReader ↓ CustomerTierProcessor implements ItemProcessor ↓ CustomerItemWriter extends AbstractItemWriter These types are what connect the application code to the Jakarta Batch runtime. Connecting the Artifacts With JSL The Java classes define the behavior, but Jakarta Batch requires explicit mapping of the reader, processor, and writer to each job. This orchestration is described in JSL: XML <?xml version="1.0" encoding="UTF-8"?> <job id="customer-segmentation" xmlns="https://jakarta.ee/xml/ns/jakartaee" version="2.0"> <step id="recalculate-customer-tiers"> <chunk item-count="20"> <reader ref="customerItemReader"/> <processor ref="customerTierProcessor"> <properties> <property name="thresholds" value="#{jobParameters['thresholds']}"/> </properties> </processor> <writer ref="customerItemWriter"/> </chunk> </step> </job> The ref values correspond directly to the names declared with @Named in the Java classes: Java @Named("customerItemReader") @Named("customerTierProcessor") @Named("customerItemWriter") The XML therefore tells the Jakarta Batch runtime: for this step, use this reader, then this processor, and finally this writer. It also maps the thresholds job parameter into the processor property. The item-count="20" sets the chunk size for this sample. Jakarta Batch coordinates reading and processing, periodically invoking the writer according to the chunk lifecycle and establishing transaction and checkpoint boundaries. The value 20 is for demonstration; real applications should tune chunk size based on processing cost, database behavior, transaction size, throughput, and recovery requirements. This structure is recommended for the article: present the class declaration first, then describe the lifecycle methods inherited from or required by Jakarta Batch. This approach helps the sample teach the API rather than simply presenting isolated methods. Conclusion Jakarta Batch is valuable because it transforms large-scale data processing into a structured execution model, eliminating the need for custom loops and ad hoc background logic. By separating reading, processing, and writing, and introducing runtime concepts such as jobs, steps, checkpoints, restartability, and chunk-oriented execution, it provides enterprise applications with a predictable approach to handling workloads involving thousands or millions of records. This allows implementations to focus on business logic, while the Batch runtime manages repetitive execution concerns, making the model easier to understand, optimize, and scale as workloads increase.
Multi-step flows are common in UX, including onboarding, checkout, account setup, approval processes, configuration wizards, and administrative tasks. These require users to move through multiple screens while continuing a consistent working state. The challenge is to keep this state active for the duration of the interaction, but not beyond. Request scope is too short, while session scope often extends longer than the business process needs. Jakarta Faces handles this with @FlowScoped, which manages state based on the lifecycle of a flow instead of a single page or the entire session. This article uses a customer segmentation application to demonstrate how a flow can guide users through configuration, preview, and confirmation, while maintaining state across each step. This approach creates a cleaner model for wizard-style UX: the scope begins when the user enters the flow, persists during navigation, and ends upon exit. Why Jakarta Faces Still Matters Jakarta Faces continues to be relevant because many enterprise applications are developed and maintained by teams with strong Java expertise. In these environments, a server-side UI framework limits context switching, keeps validation and navigation close to the application model, and allows teams to reuse the same language, dependency injection model, and enterprise APIs throughout the stack. Architecturally, if the team is proficient in Java and the application is form-driven, workflow-oriented, or back-office focused, introducing a separate frontend stack does not necessarily offer an advantage. Component libraries such as PrimeFaces further support this approach. Rather than building tables, dialogs, forms, charts, wizards, and validation from scratch, teams can use reusable components within the Jakarta EE programming model. This can accelerate delivery and lessen the need for custom frontend infrastructure. While the decision should be based on product and team context, for Java-focused enterprise teams, Jakarta Faces is a pragmatic architectural choice, not just a legacy option. A few publicly documented examples of organizations that have used Jakarta EE/JSF or PrimeFaces include: NASAWalmart LabsLufthansaRakutenCommerzbankUnited NationsPenn State UniversityHarvard UniversityTelefonicaBig LotsComfortel (telecommunications)Various commercial banks and financial institutions Overview of Jakarta Faces Scopes Jakarta Faces applications maintain managed bean state for varying durations. Choosing the right scope is an architectural decision and must match the user interaction's lifetime. Some state is limited to a single HTTP request, a single page, a multi-step flow, or the entire user session. @RequestScoped is the shortest-lived scope. It creates a bean instance for each HTTP request and discards it when the request completes. It suits stateless actions, simple submissions, and operations that don't need to continue across navigation or Ajax interactions.@ViewScoped retains the bean while the user stays on the same Faces view. It is ideal for pages with forms, tables, filtering, pagination, dialogs, or Ajax interactions that update the same page multiple times. The state is discarded when the user navigates to a different view.@FlowScoped is intended for business interactions spanning multiple views, such as checkout, onboarding, configuration wizards, approval workflows, or customer segmentation. The bean remains active throughout the flow and is destroyed when the flow ends. Its lifecycle falls between view scope and session scope.@SessionScoped maintains state for the entire user session. It suits information that remains across multiple pages, such as user preferences or session-level context. Avoid using it for temporary workflow state, as this can unnecessarily extend the state’s lifetime.@ApplicationScoped has the broadest lifetime, sharing a single bean instance across the entire application and all users. It suits shared services, caches, configuration, or application-wide state, but not per-user or per-flow data unless explicitly designed for sharing and thread safety. Building a Multi-Step Experience With @FlowScoped This article focuses on Jakarta Faces flow, which models user engagements spanning multiple pages but shorter than a full HTTP session. In the customer-segmentation example, the administrator configures thresholds, previews their impact, reviews the configuration, and starts the operation. These steps form a single business process and should share the same state. Jakarta Faces represents this process with a flow definition and a flow-scoped managed bean. In this project, the flow resides in the customer-segmentation directory: reStructuredText src/main/webapp/ └── customer-segmentation/ ├── customer-segmentation-flow.xml ├── configure.xhtml ├── preview.xhtml └── review.xhtml Aligning the directory, flow identifier, and bean name clarifies their relationship. Here, the flow stays named customer-segmentation, defined in customer-segmentation/customer-segmentation-flow.xml, and the Java bean uses @FlowScoped("customer-segmentation"). The value passed to @FlowScoped ties the bean’s lifecycle to the corresponding Faces flow. The XML file defines the flow’s structure and navigation boundaries, but does not store business state. In this example, configure is the starting point, followed by preview and review. Each view has an identifier and references its corresponding XHTML document: XML <flow-definition id="customer-segmentation"> <start-node>configure</start-node> <view id="configure"> <vdl-document> /customer-segmentation/configure.xhtml </vdl-document> </view> <view id="preview"> <vdl-document> /customer-segmentation/preview.xhtml </vdl-document> </view> <view id="review"> <vdl-document> /customer-segmentation/review.xhtml </vdl-document> </view> <flow-return id="home"> <from-outcome>/index?faces-redirect=true</from-outcome> </flow-return> </flow-definition> The start-node specifies where the interaction begins. The <view> elements define the flow’s pages, and <flow-return> determines how the application exits. Returning the outcome home ends the flow, redirects the user to the dashboard, and discards the flow-scoped state. Navigation within the flow preserves the state, while exiting ends the conversation. On the Java side, CustomerSegmentationFlow manages the conversation state as both a named Faces bean and a flow-scoped bean: Java @Named @FlowScoped("customer-segmentation") public class CustomerSegmentationFlow implements Serializable { @Inject private CustomerSegmentationFlowService flowService; private CustomerSegmentationFlowState state; @PostConstruct public void initialize() { state = flowService.initializeState(); } // ... } @Named makes the bean accessible to Faces pages through Expression Language, while @FlowScoped("customer-segmentation") assigns its lifecycle to the specific flow. The state created during @PostConstruct remains across requests and page transitions during the flow, so CustomerSegmentationFlowState remains available throughout the interaction. This distinction sets @FlowScoped apart from @ViewScoped. With view scope, moving from configure.xhtml to preview.xhtml starts a new conversation. With flow scope, both views remain part of the same business interaction. The scope follows the conversation, not individual pages. Navigation methods on the bean then return outcomes that correspond to nodes defined by the flow: Java public String preview() { flowService.preview(state); return "preview"; } public String review() { return "review"; } Returning "preview" moves the user to the preview view in the flow definition, while "review" advances to the next view. Since both destinations are within customer-segmentation, the same flow-scoped bean and its state remain active. The final action demonstrates the other side of the lifecycle: Java public String execute() { long executionId = flowService.start(state); // message handling omitted return "home"; } Home is not a page within the wizard. It corresponds to the <flow-return id="home"> element in the XML definition. When this outcome occurs, Faces exits the flow, redirects to the dashboard, and discards the associated state. @FlowScoped is ideal for processes such as checkout, onboarding, registration, approval, configuration, and administrative wizards. It provides a scope broader than a single page but narrower than a session. Instead of using @SessionScoped for temporary workflow data, the state exists only for the workflow's duration. This article covers the Faces flow: how pages are connected, how state persists between them, and how entering and exiting the flow manages the Java bean’s lifecycle. The full application also includes MongoDB persistence, dashboard services, preview calculations, and Jakarta Batch processing. The complete source code is available at https://github.com/soujava/mongodb-jakarta-batch. Here, the emphasis stays on the user interaction represented by Configure → Preview → Review → Exit. Conclusion @FlowScoped provides Jakarta Faces with an effective way to model multi-step user experiences as a single business conversation. Rather than storing temporary workflow state in @SessionScoped or reconstructing it between views, the flow maintains state only while the user progresses through the defined steps and releases it when the flow concludes. For Java-focused enterprise teams, this approach simplifies wizards, onboarding, approvals, checkout flows, and configuration processes by keeping navigation, state, and lifecycle consistent with the user experience.
Large agent prompts often begin as a practical shortcut: policies, domain rules, tool descriptions, examples, recovery procedures, and integration notes are placed in one system message so every capability is always available. That approach stops scaling once an agent accumulates dozens of tools and specialized workflows. Tool definitions and instructions consume context on every turn, irrelevant material competes with task-relevant material, and each integration enlarges a shared prompt that becomes harder to test and version. Current platform guidance increasingly converges on a different model: expose compact capability metadata first, load detailed instructions only after relevance is established, and execute specialized logic inside controlled tool or sandbox boundaries. Anthropic describes this as progressive disclosure for Agent Skills, while OpenAI supports both Skills and deferred tool discovery. Context Should Be Earned, Not Prepaid Progressive disclosure treats context as a runtime resource rather than a static configuration file. A skill registry initially contributes only descriptors such as name, purpose, version, input shape, side-effect class, and required capabilities. When intent matches a descriptor, the runtime loads the skill’s main instructions. Deeper references, scripts, templates, or schemas remain outside active context until needed. Anthropic’s skill model formalizes the same layering: metadata is the first disclosure level, the full SKILL.md is the second, and linked supporting files form later levels. OpenAI’s Skills documentation similarly exposes name and description during discovery, then lets the model read full instructions and supporting files after selection. A minimal runtime contract can keep selection separate from execution: Java @Skill(id = "invoice.reconcile", version = "3", risk = "read") public SkillResult invoke(SkillRequest request) { SkillDescriptor descriptor = registry.describe(request.skillId()); SkillPackage skill = registry.load(descriptor.id(), descriptor.version()); policy.authorize(request.principal(), descriptor, request.arguments()); return sandbox.execute(skill, request.arguments(), request.deadline()); } The important boundary is the order of operations. describe is metadata-oriented; load materializes selected instructions and resources; authorize evaluates the proposed operation independently of model reasoning; sandbox.execute provides an execution boundary. Skill discovery therefore does not imply permission, and packages can evolve independently while the core agent prompt stays small. The motivation is not merely context-window capacity. Anthropic’s current context guidance notes that system prompts, messages, tool results, and tool definitions all consume context, and that larger context can degrade recall and accuracy as token counts rise. OpenAI’s tool-search interface consequently allows selected function definitions to be deferred until discovery instead of exposing every definition eagerly. Discovery Is a Protocol Concern Once skills become modular, capability negotiation becomes as important as prompt composition. A descriptor should state what a skill needs before activation: structured output, file access, network access, long-running execution, approval support, or a protocol version. The runtime should intersect those requirements with host support and policy. Selection can then fail early instead of allowing an incompatible skill into the reasoning loop. Java public NegotiatedCapabilities negotiate( AgentCapabilities agent, SkillDescriptor skill, PolicyScope scope) { return agent.intersect(skill.requiredCapabilities()) .restrictTo(scope.allowedCapabilities()) .require(skill.minimumProtocolVersion()); } MCP provides a useful reference model even when MCP is not used directly. In the 2026-07-28 specification, server/discover returns supported versions and server capabilities, while requests carry protocol version and client capability metadata. The same release adds ttlMs and cacheScope to cacheable discovery results and supports change notifications for tool lists. These mechanisms matter because production capability catalogs are dynamic: tools can disappear because of permissions, outages, tenancy, or deployments. Cached discovery therefore needs explicit freshness semantics. A practical registry can keep a small cacheable index of descriptors and version pointers while storing full skill bodies separately. Version pinning prevents an active run from silently switching behavior mid-task. Long-lived business state should also remain outside the prompt as structured run state, artifact references, or domain records. OpenAI’s Agents documentation similarly treats history, continuation identifiers, interruptions, and resumable state as explicit runtime surfaces rather than one text transcript. Execution Boundaries Matter More Than Prompt Boundaries Progressive disclosure reduces exposure, but it does not make a skill trustworthy. Skill instructions can contain executable scripts, tool calls, file references, and untrusted text. OpenAI warns that skills can introduce prompt-injection-driven data exfiltration and recommends review before exposure; Anthropic’s programmatic tool-calling guidance distinguishes unsafe local execution from sandboxed execution with restrictions such as disabled network egress. The safer design treats model output as a proposal. Authorization should be enforced beside the side effect, using independently computed identity, tenant, scope, destination, and argument constraints. Read-only skills can receive broader automatic execution, while write, shell, credential, or external-network skills can require approval. OpenAI’s guardrail guidance makes the same boundary explicit: tool arguments and results can be checked at the tool boundary, and sensitive side effects can pause for human approval. Fallback behavior also belongs in the contract rather than in a vague prompt instruction: Java @SkillFallback(forSkill = "customer.profile") private SkillResult fallback(ProfileRequest request, SkillException ex) { if (ex.retryable()) { return SkillResult.retry("profile-cache", request.customerId()); } return SkillResult.partial("profile unavailable", ex.errorCode()); } This distinguishes recoverable infrastructure failure from semantic failure. A fallback may choose a cached or lower-fidelity capability, but it should preserve the original authorization scope and return structured provenance indicating degraded execution. Silent fallback to a more privileged tool is an anti-pattern because availability logic then becomes privilege escalation. Production Behavior Needs Evidence Progressive disclosure introduces a measurable trade-off. Smaller active context can reduce token usage and model distraction, but discovery, loading, and sandbox startup add latency. Anthropic reports that programmatic tool calling reduced billed input tokens by about 38% on a 75-tool benchmark, yet cost about 8% more on a benchmark dominated by one or two sequential tool calls. The broader implication is that eager loading remains reasonable for a tiny stable core, while specialized or heavy capabilities benefit more from on-demand activation. Testing should cover more than final answer quality. Skill-selection tests should verify relevant activation and rejection of near-neighbor skills. Contract tests should validate schemas, capability requirements, version compatibility, timeouts, fallback semantics, and policy denial. Sandbox tests should exercise filesystem and network boundaries. End-to-end evaluations should score complete traces, including tool choice, routing, and policy behavior; OpenAI’s evaluation guidance supports trace grading across model calls, tool calls, guardrails, and handoffs. Observability should expose the same lifecycle as the runtime. Useful spans include discovery, descriptor match, package load, authorization, invocation, fallback, and completion, with skill ID, resolved version, latency, token counts, sandbox identity, policy decision, and outcome attached as structured attributes. Sensitive arguments should be redacted. OpenAI tracing already records agent and tool spans, durations, errors, arguments, results, and token usage, providing a concrete precedent for this level of visibility. Incremental rollout is safer than replacing a giant prompt in one release. Existing prompt logic can first run beside a metadata registry in shadow mode, producing selection decisions without executing skills. Read-only skills can then move behind feature flags, followed by canary traffic for side-effecting skills with approval enforced. Versioned bundles and explicit registry pointers make rollback deterministic. As evidence accumulates, stable instructions can leave the monolithic prompt and become independently deployable capabilities. An extensible agent does not need an ever-growing prompt; it needs a small stable core, a discoverable capability surface, explicit negotiation, controlled execution, durable external state, and observable contracts. Progressive disclosure turns agent growth from prompt accumulation into modular software composition. The resulting system spends context only when a capability is relevant, keeps authorization outside model judgment, isolates risky execution, and permits skills to be versioned, tested, rolled out, and replaced independently. That shift is the practical path from a brittle all-knowing prompt toward an agent platform that can expand without making every task carry the weight of every capability.
When many developers think about recommendation engines, they think of machine learning: collaborative filtering models, matrix factorization, embedding vectors, and training pipelines. What surprises many people is that you can build a genuinely useful recommendation system with nothing more than a graph database and several Cypher queries. No scikit-learn, no TensorFlow, no model training. Just the natural structure of the data doing the work. In this article, we'll build a product recommendation engine on top of Neo4j Aura using two Jupyter notebooks. The first generates a realistic synthetic dataset and loads it into Aura. The second runs four recommendation queries directly in Cypher and visualizes the results with Plotly. Everything runs locally in a Python virtual environment against a free cloud Neo4j instance. The full source code is available on GitHub. Why Graphs Are a Natural Fit for Recommendations The core intuition behind most recommendation approaches is relationship: this customer bought that product, those products appear together in the same order, this product shares attributes with that one. In a relational database, capturing these relationships means multiple self-joins across large tables. A query like "find products bought by customers who also bought what this customer bought" quickly becomes difficult to write and expensive to execute at scale. In a graph, that same question is a traversal. We follow edges from a customer to the products they purchased, hop across to other customers who share those products, and collect what else those customers bought. The query is short, the intent is clear, and the graph engine is optimized for exactly this kind of path-following work. Prerequisites AuraDB is Neo4j's fully managed cloud database. A free tier is available with no credit card required. Sign up at Get Started for Free.Create a new AuraDB Free instance.When the instance is created, download or note the credentials — the connection URI, username, and password.Once the instance is running, open the Query tab and connect to the instance.Confirm it's empty with MATCH (n) RETURN count(n) which should return 0 A virtual environment is highly recommended. For example: Shell python3 -m venv ~/recommendation-engine-env source ~/recommendation-engine-env/bin/activate Before starting Jupyter, export the connection details as environment variables in your shell: Shell export NEO4J_URI="neo4j+s://xxxx.databases.neo4j.io" export NEO4J_USERNAME="your_username_here" export NEO4J_PASSWORD="your_password_here" The Graph Model Before we write any code, let's define the graph. We have four node types and three relationship types. Nodes Customer – id, name, email, city, country.Product – id, name, description, price.Category – name (e.g., Electronics, Clothing, Books).Tag – name (e.g. "wireless", "eco-friendly", "premium"). Relationships (:Customer)-[:PURCHASED {order_id, quantity, order_date}]->(:Product) — order metadata lives on the relationship rather than a separate Order node, which keeps our Cypher clean.(:Product)-[:BELONGS_TO]->(:Category)(:Product)-[:TAGGED_WITH]->(:Tag) The decision to put order_id, quantity and order_date on the PURCHASED relationship is worth discussing. It means a single customer can have multiple PURCHASED relationships to the same product (each with a different order_id) and we can group by order_id to find products that appeared together in the same basket — which is exactly what our co-purchase query needs. Figure 1 illustrates exactly this point, as we have a customer, two products, and the same order_id. Figure 1. Shared order_id enables co-purchase queries Notebook 1: Data Generation and Loading Rather than sourcing an external dataset, we'll generate synthetic data using Faker. This keeps the notebook fully self-contained, and readers can run it as-is without downloading anything. We'll generate 2,000 customers, 500 products across 15 categories, and 20,000 orders. Each order is a basket of several products sharing the same order_id — this is the key design decision that makes the frequently-bought-together query work. With an average basket of 3 products, we end up with around 60,000 PURCHASED relationships in the graph. Realistic Product Names Faker's default catch_phrase() method produces output like "Proactive exuding encoding" — readable enough for a demo but not really useful in an article. Instead, we define a PRODUCT_VOCAB dictionary keyed by category, each containing lists of adjectives, nouns, use cases, and benefit statements. A product name is then a simple combination, as follows: Python def make_product_name(category): vocab = PRODUCT_VOCAB[category] adj = random.choice(vocab["adjectives"]) noun = random.choice(vocab["nouns"]) return f"{adj} {noun}" def make_product_description(category, name): vocab = PRODUCT_VOCAB[category] use_case = random.choice(vocab["use_cases"]) benefit = random.choice(vocab["benefits"]) return f"The {name} is designed for {use_case}. {benefit}." This gives us names like "Wireless Noise-Canceling Earbuds," "Organic Ground Coffee" and "Ergonomic Lumbar Support Cushion" — realistic enough to make the recommendation output meaningful. Basket-Based Order Generation Each order picks a random customer, generates a unique order_id, and samples several products into a basket. We then flatten the basket into individual order lines, each carrying the shared order_id: Python orders = [] for _ in range(NUM_ORDERS): order_id = str(uuid.uuid4()) customer = random.choice(customers) order_date = (start_date + timedelta(days=random.randint(0, 730))).strftime("%Y-%m-%d") basket = random.sample(products, k=random.randint(2, 4)) for product in basket: orders.append({ "order_id": order_id, "customer_id": customer["id"], "product_id": product["id"], "quantity": random.randint(1, 5), "order_date": order_date }) Loading Into Aura Data loading is in batches of 100 using MERGE statements. To show progress during the load, we'll use tqdm as ~60,000 order lines can take several minutes, and the progress bars make it easy to see what's happening: Python with driver.session() as session: customer_batches = range(0, len(customers), BATCH_SIZE) for i in tqdm(customer_batches, desc="Loading customers", unit="batch", colour="#1f77b4"): session.execute_write(load_customers, customers[i:i+BATCH_SIZE]) product_batches = range(0, len(products), BATCH_SIZE) for i in tqdm(product_batches, desc="Loading products ", unit="batch", colour="#1f77b4"): session.execute_write(load_products, products[i:i+BATCH_SIZE]) for product_id, tags in tqdm(product_tags.items(), desc="Loading tags ", unit="product", colour="#1f77b4"): session.execute_write(load_tags, product_id, tags) order_batches = range(0, len(orders), BATCH_SIZE) for i in tqdm(order_batches, desc="Loading orders ", unit="batch", colour="#1f77b4"): session.execute_write(load_orders, orders[i:i+BATCH_SIZE]) A verification query at the end confirms the counts. The Four Recommendation Queries Notebook 2 runs four Cypher queries against the loaded graph, each implementing a different recommendation strategy. Before running any query, we fetch a stable seed customer, product, and category: Python with driver.session() as session: customer = session.run(""" MATCH (c:Customer) RETURN c.id AS customer_id, c.name AS customer_name ORDER BY c.name ASC LIMIT 1 """).single() product = session.run(""" MATCH (p:Product)<-[r:PURCHASED]-() RETURN p.id AS product_id, p.name AS product_name, count(r) AS order_count ORDER BY order_count DESC LIMIT 1 """).single() top_cat = session.run(""" MATCH (p:Product)-[:BELONGS_TO]->(cat:Category) RETURN cat.name AS category, count(p) AS total ORDER BY total DESC LIMIT 1 """).single() We pick the alphabetically first customer for consistency, the most-purchased product to ensure co-purchase data exists, and the category with the most products for the trending query. This makes the notebook reproducible across runs. Query 1: Collaborative Filtering The classic "customers who bought this also bought" approach. We find customers who share at least one purchased product with the seed customer, then collect what else those customers bought — excluding anything the seed customer already purchased. Python def collaborative_filtering(tx, customer_id, limit=5): result = tx.run(""" MATCH (target:Customer {id: $customer_id})-[:PURCHASED]->(p:Product) <-[:PURCHASED]-(other:Customer)-[:PURCHASED]->(rec:Product) WHERE NOT (target)-[:PURCHASED]->(rec) RETURN rec.id AS id, rec.name AS product, rec.price AS price, count(other) AS score ORDER BY score DESC, id ASC LIMIT $limit """, customer_id=customer_id, limit=limit) return result.data() The score is the number of other customers whose purchasing overlap with our target customer also led them to buy the recommended product. A higher score means more customers in the overlap group bought it, making it a stronger signal. In Cypher, the traversal reads almost like the description: start at the target customer, follow PURCHASED edges to products, hop to other customers who bought the same products, then follow their PURCHASED edges to new products. Query 2: Frequently Bought Together This query finds products that appeared in the same order as the seed product. The key is matching on order_id across two PURCHASED relationships from the same customer: Python def frequently_bought_together(tx, product_id, limit=5): result = tx.run(""" MATCH (p:Product {id: $product_id})<-[r1:PURCHASED]-(c:Customer) -[r2:PURCHASED]->(other:Product) WHERE r1.order_id = r2.order_id AND other.id <> $product_id RETURN other.id AS id, other.name AS product, other.price AS price, count(c) AS frequency ORDER BY frequency DESC, id ASC LIMIT $limit """, product_id=product_id, limit=limit) return result.data() The WHERE r1.order_id = r2.order_id clause is what makes this work. It constrains the traversal to only consider cases where both products were part of the same order, not just bought by the same customer at different times. frequency counts how many distinct customers placed an order containing both products together. Query 3: Content-Based Filtering Rather than looking at purchase behavior, this query finds products similar to the seed product based on shared tags. The more tags two products have in common, the more similar they are: Python def content_based(tx, product_id, limit=5): result = tx.run(""" MATCH (p:Product {id: $product_id})-[:TAGGED_WITH]->(t:Tag) <-[:TAGGED_WITH]-(rec:Product) WHERE rec.id <> $product_id RETURN rec.id AS id, rec.name AS product, rec.price AS price, count(t) AS shared_tags ORDER BY shared_tags DESC, id ASC LIMIT $limit """, product_id=product_id, limit=limit) return result.data() The traversal goes outward from the seed product through its tags, then back inward to any other product that shares those same tags. count(t) gives the number of shared tags, which serves as a simple but effective similarity score. This approach works without any purchase history, making it useful for recommending products to new customers or for newly listed products with no order data yet. Query 4: Trending in Category This query finds the most purchased products in the top category within a fixed date window. In our case, this is from 2024-10-01 onwards: Python def trending_in_category(tx, category_name, cutoff="2024-10-01", limit=5): result = tx.run(""" MATCH (p:Product)-[:BELONGS_TO]->(cat:Category {name: $category_name}) MATCH (:Customer)-[r:PURCHASED]->(p) WHERE date(r.order_date) >= date($cutoff) RETURN p.id AS id, p.name AS product, p.price AS price, count(r) AS purchases ORDER BY purchases DESC, id ASC LIMIT $limit """, category_name=category_name, cutoff=cutoff, limit=limit) return result.data() We use date() conversion on the stored string order_date to enable date comparison. count(r) counts individual PURCHASED relationships rather than distinct customers, so a customer who bought the same product multiple times within the window is counted each time — reflecting genuine demand volume rather than unique buyer count. Notebook 2: Results Each query outputs a table followed by a Plotly horizontal bar chart. Here are the results for our seed data. Collaborative Filtering Figure 2 returns five products. The top recommendation is An Introduction to Public Speaking, driven by the number of customers whose purchasing overlap with Aaron Boyd also led them to buy it. Heavy-Duty Cable Management Box and Educational Coding Robot follow closely, showing that the overlap group bought broadly across categories rather than clustering in one area. Figure 2. Collaborative filtering Frequently Bought Together Figure 3 shows products co-purchased with the Durable Grooming Brush in the same order basket. The top results — Waterproof Hammock and Natural Body Lotion at frequency 4, followed by Adjustable Lumbar Support Cushion, Sugar-Free Collagen Powder and Slim-Fit Hiking Vest at frequency 3 — show which products most commonly appeared alongside the seed product in the same order. The cross-category spread here (Beauty, Outdoor, Clothing, Health, Office) is a feature of random synthetic data; in a real system, we'd expect more category clustering. Figure 3. Frequently bought together Content-Based Filtering Figure 4 finds products sharing the most tags with the seed product. All five results share 2 tags with the Durable Grooming Brush — Smart Mechanical Keyboard, Waterproof Toiletry Bag, Ergonomic Whiteboard, Cold-Pressed Hot Sauce, and Durable Dumbbell Pair. The cross-category reach (Sports, Food & Drink, Office, Travel, Electronics) illustrates the tag graph doing its job: shared attributes like "durable" or "waterproof" create similarity links that cross category boundaries, which is useful for surface-level discovery recommendations. Figure 4. Content-based filtering Trending in Category Figure 5 shows the top 5 products in Toys — the category with the most products in our graph — with purchase counts from 2024-10-01 onwards. Battery-Free Coding Robot leads, followed by Battery-Free Building Blocks Set, Interactive Remote Control Car, Wooden Magnetic Drawing Board, and Creative Puzzle Game. The scores are tight here, which makes sense because within a single category over a fixed time window, popular products tend to cluster around similar purchase volumes. Figure 5. Trending Summary We've built a working product recommendation engine using nothing but Neo4j, Cypher, and a few Python libraries. No ML framework, no training data, no model deployment. The four queries cover the most common recommendation patterns in production systems: Collaborative filteringCo-purchase analysisContent similarityTrending detection The graph model is the foundation that makes this possible. Storing orders as relationships with properties means co-purchase queries are a natural traversal rather than a complex join. Adding tags as nodes means similarity queries are just path-matching. Because everything lives in the same graph, we can also combine these approaches. For example, filtering collaborative filtering results by tag similarity using a single extended Cypher query. The full source code is available on GitHub.
Editor’s Note: The following is an article written for and published in DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale. Every engineering organization that I have worked with eventually faces the same issue, which is that each team ships services differently. One team used Helm, another wrote raw manifests, and a third would have built a custom Bash script. As these different approaches accumulate, the supporting deployment steps often end up scattered across multiple Wiki pages that quickly go stale. New engineers then spend their first two weeks copying configuration values from an old repository and hoping they still work. A golden path fixes this without turning the platform team into a gatekeeper. It provides users with a standardized workflow for the shortest and most obvious route from a fresh repo to a production workload. This guide walks you through designing a minimum viable golden path, where guardrails belong, and how to keep it useful after v1. Choose the First Golden Path Start with one workflow to standardize first; the strongest candidate is usually the workflow your teams ship most often, or one that teams experience the most friction with. In many organizations, that workflow is a stateless HTTP service exposing a REST or gRPC API endpoint, deployed to Kubernetes and owned by one application team. For this walkthrough, we will use orders-api, a stateless HTTP service on Kubernetes, as our reference throughout this article. The intended users are application developers, not platform engineers — those who create the golden path itself. The path starts with a create-service command in a CLI or a form in an internal developer portal. It should end when the service is running in production with logs, metrics, ownership, and on-call rotation attached. Keep the first version deliberately narrow. A workload that needs GPU nodes, a queue-driven scaling model, or a stateful sidecar can wait. Trying to capture every exception at the beginning turns a practical delivery path into a long platform program. A golden path’s success criteria are qualitative, not quantitative. Analyze the first release by user adoption and experience. Are teams using standardized workflows instead of copying an old repository? Can a new engineer understand the end-to-end deployment process without asking around? Are on-call handoffs easier because services have the same operational shape? The answers to these questions matter more than looking at any adoption numbers displayed on a dashboard in the first few months. Define What the Path Standardizes A golden path is a curated set of decisions that are made once and reused consistently across services: The workload template should provide a Dockerfile, fully maintained base image, Kubernetes manifests, probes, resource requests and limits, a Pod Disruption Budget (PDB), autoscaling defaults, and consistent labels.The delivery pipeline should build, test, scan, sign, and publish the image.The platform defaults should include namespace rules, quotas, network policies, ingress, TLS, logging, metrics, tracing, and basic alerts. The path should not own product decisions; teams will still choose their language, framework, business logic, schema, feature flags, test strategy, and service-specific objectives. This boundary is very important. If we over-standardize, developers will work around the platform, and if we under-standardize, every instance will start with a different set of commands and dashboards. Also make sure the path is easy to find. One internal documentation page, one command, and one entry in the developer portal are enough. If a developer has to ask which template to use, the path has already failed and created friction. The table below shows the differences between shared standards the path owns and decisions each service team owns. Shared Standards vs. Team-Owned Decisions shared standard team decision Dockerfile, base image, patching cadence Language and framework choiceDeployment manifests, probes, resource requests/limits, PDB, Horizontal Pod Autoscaler Business logic, schema, feature flags Build, test, scan, sign, and publish pipeline Test suites specific to the service Namespaces, quotas, network policies, ingress, and TLS defaults Non-standard scaling (queue-driven consumers, GPU jobs) Logging, metrics, tracing, and alerting defaults Business-specific dashboards and SLOs Turn Common Requests Into Self-Service Actions Once the path is created and available to users, review the top 10 tickets your platform team receives. Look for repeated requests such as creating namespaces, adding a database, registering a DNS name, rotating a secret, or creating another environment. These are all good candidates because the desired outcome is already understood, and the steps are mostly predictable. For the Orders API golden path, the platform team can provide the following self-service actions and apply guardrails based on the risk from each change: Fully automated. These actions are reversible and have a limited blast radius. Creating a development namespace for orders-api, spinning up a preview environment on a PR, or rotating a non-production secret happens on demand without a human involved to review.Light review. Actions that change cost, security exposure, or shared infrastructure should require a light review. Provisioning production Postgres for orders-api opens a pre-filled change request that needs one approval. A new public DNS record on a shared domain is reviewed through a one-click approval on a pre-filled PR.Approval mechanism. Every self-service action generates a PR against a config repo, pre-fills the values, tags the reviewer, and merges on approval. The change flows through the same pipeline as code, and every action leaves an audit trail because it’s a git commit. The self-service interface should offer supported choices instead of exposing raw cloud APIs. For example, allowing every team to choose any PostgreSQL version, instance class, or backup schedule can leave the platform team operating 30 different database configurations. A better approach is to provide a small, opinionated set of options such as small, medium, and large. This gives developers enough flexibility while keeping the operational model understandable. For our Orders API, the developer-facing configuration can stay small: YAML # svc.yaml name: orders-api owner: team-orders tier: standard # small | standard | high runtime: http dependencies: - kind: postgres size: small # opinionated preset, not raw config on_call: orders-oncall The configuration captures the developer’s intent, while the golden path translates each request into an approved action with the right guardrail and a clear record of what happened. The table below shows how this works for the Orders API. Orders API Self-Service Actions, Guardrails, and Evidence Step Self-Service Action Guardrail Evidence Create service Run svc new via CLI or submit a portal form Template pinned to current version; namespace quotas applied Repository created with owner metadata; entry in service catalog Add dependency Pick from opinionated list (small/medium/large DB) One-click PR review for prod-tier resources Merged PR against config repo with reviewer name Deploy to prod Merge to main triggers promotion Progressive rollout with auto-rollback on error/latency signals Deployment record with canary metrics and rollback status Rotate secret Run svc rotate-secret New version issued; old version revoked after grace window Audit log entry linked to requester Create a Consistent Path From Code to Deployment Every service on the golden path should move through the same basic stages: pull request → merge to main → staging → production. The exact tooling can vary, but the meaning of each stage should not. At the PR stage, CI runs unit tests, linting, the container build, and security checks. Produce an immutable image tagged with the commit identifier, but do not deploy it to production.On merge to main, the same image is promoted to staging automatically. Rebuilding at each stage creates uncertainty because the artifact tested is no longer guaranteed to be the artifact released. Run integration and smoke tests in this stage.Promoting the image to production reveals the delivery guardrails. Start with a small percentage of traffic (5-10%), monitor health signals, and continue increasing traffic to 25%, then 100%. Roll back automatically when error rate, latency, or probe failures cross agreed thresholds. A developer should not have to recreate this logic in every repository — it should be baked into the deployment tooling. A failed orders-api canary would look like this end to end: The pipeline promotes the new image to 5% of production pods.The error rate for the /orders endpoint rises sharply during the observation window.The deployment controller restores the previous image and drains the new pods based on the rollback threshold.The pipeline posts a message in the orders-oncall service channel with a link to the failing dashboard and offending commit identifier (SHA).An incident record is created automatically only when rollback fails, or the service remains unhealthy. Teams may skip a stage for a documented case (e.g., configuration-only change), but the exception should be an explicit setting with an owner, not an informal workaround. Plain Text # pipeline stages (pseudo) on_pr: [test, lint, build, scan, sign] on_merge: [promote_to_staging, integration-tests] on_green: [canary-5, wait-signals, canary-25, wait-signals, full-rollout] On_regress: [auto-rollback, notify-oncall, record-failure, open-incident] Observability and Day-1 Operational Defaults Even if its pods are running, a service is not ready until the owning team can determine whether it is healthy and knows what action to take when it is not. The golden path should therefore create the minimum operational surface at the same time as the service. The template includes the following list on day one: Structured logs to the central log store, with request ID and trace identifiersRequest rate, error rate, latency percentiles, and saturation metricsDistributed traces with a platform-managed sampling defaultA standard dashboard created from the service nameAlerts for high errors, high latency, restart loops, and resource pressureLiveness and readiness checks connected to a health endpoint Ownership should also be captured during service creation. Ask for the team, on-call rotation, and support channel, then reuse those values in alert routing, the service catalog, and the runbook. Generate a simple runbook with sections dedicated to common failures such as stalled deployments, elevated errors, and pod eviction. A partially completed runbook with a familiar structure is far more useful than a blank page, and consistency here pays off during an incident. Keep the Golden Path Useful Over Time Exceptions are inevitable, so record the failure reason, owner, and expiry date rather than letting the exception become a permanent member. At review time, either the service returns to the path or the platform team decides the pattern is common enough to support. Treat templates and defaults like product code: review changes, version them, and provide a propagation method. When a base image or manifest default changes, open a change against each service instead of relying on teams to notice a document update. Silent drift is one of the fastest ways to lose developer trust in the path. Track a small set of signals such as the time from service creation to first production deployment, template version distribution, open exceptions, and the percentage of new services created through the path. Pair those numbers with developer feedback. A slow step that teams repeatedly bypass tells you where the next path improvement belongs. A new template version without a propagation plan becomes a fork. Extend the path when a pattern is used by three or more teams, but keep it narrow while it is still one team’s edge case. Plain Text # template bump propagation (pseudo) on template_release(new_version): for svc in services_on_path(): open_pr(svc, bump_template = new_version, auto_merge = svc.opts.auto_bump, reviewer = svc.owner) Making the Golden Path Useful in Practice A golden path succeeds when it is easier to follow than to work around. Start with one common workflow, standardize what is shared, and leave product choices with the service team. Make routine actions self-service, place checks in the delivery flow, and include observability from the first deployment. Usage signals can then inform future improvements to the path. A small path that ships, earns trust, and changes steadily will have a greater impact on engineering speed than a broad platform program that remains unfinished. Resources: CNCF TAG App DeliveryOpenTelemetry General Semantic ConventionsKubernetes Pod Security StandardsBackstage Software Templates“Building a CI/CD Pipeline With Kubernetes” by Naga Santhosh Reddy VootukuriKubernetes Security Essentials, DZone Refcard by Yitaek HwangPlatform Engineering Essentials, DZone Refcard by Apostolos Giannakidis This is an excerpt from DZone’s 2026 Trend Report, Cloud-Native Foundations: Kubernetes, Platform Engineering, and Distributed Operations at Scale.Read the Free Report
Unit testing business logic often requires surprisingly little business data. Suppose we want to test a program that loads an order, calculates its total, and rejects it when the amount exceeds a limit. The decision we want to verify is simple: The program loads the specified order.It asks for the order’s total.If the total is too high, it does not approve the order.It returns FALSE. Yet a conventional Java unit test may have to construct an Order. That order may require a customer, line items, currencies, prices, tax information, identifiers, and other objects that have nothing to do with the decision being tested. Builders, fixtures and mocking frameworks reduce the typing, but they do not eliminate the underlying problem: the test must participate in the internal representation of the domain model. BUBAS takes a different approach. Domain objects are opaque to a BUBAS program. Because the program cannot inspect them, a unit test does not need to construct them. It needs only a token. The Business Program Consider this BUBAS program: SQL PROGRAM ApproveOrder(orderId INTEGER, limit DECIMAL) RETURNS BOOLEAN DECLARE purchase Order DECLARE total DECIMAL purchase = LOAD_ORDER(orderId) IF NOT ORDER_WAS_FOUND(purchase) THEN LOG_EVENT "ERROR", "no such order: " + orderId RETURN FALSE END IF total = ORDER_TOTAL(purchase) IF total > limit THEN LOG_EVENT "INFO", "over limit: " + total RETURN FALSE END IF APPROVE purchase RETURN TRUE END. Order is a Java domain type registered by the application embedding BUBAS. The program can store an Order in a variable and pass it to operations that accept an Order, but it cannot access its fields or invoke its methods. There is no expression such as: SQL purchase.customer.account.balance If the program needs information about an order, the application must expose an operation for obtaining it: SQL total = ORDER_TOTAL(purchase) This restriction is primarily an encapsulation mechanism. The business program depends on the vocabulary of its domain rather than on the internal structure of Java objects. It also has an important consequence for testing. Replace the Object With Identity Here is a BUNIT test for the over-limit case: Gherkin PROGRAM OverLimitIsRejected "LOAD_ORDER" WITH ARGS(42) RETURNS "o1" "ORDER_TOTAL" WITH ARGS("o1") RETURNS 1500.00 "APPROVE _" IS MOCKED ARGUMENT "orderId" IS 42 ARGUMENT "limit" IS 1000.00 RUN RESULT IS FALSE "APPROVE _" WAS NOT CALLED END. The string "o1" is not an order serialized as text. It does not contain an order number, a total or any other property. It is a test token representing one opaque Order. The first mock says: Plain Text "LOAD_ORDER" WITH ARGS(42) RETURNS "o1" When the program calls LOAD_ORDER(42), BUNIT returns the token "o1" in place of the real Java object. The program stores it in purchase. Later it calls: Java ORDER_TOTAL(purchase) The second mock recognizes that same token and returns 1500.00. The program cannot tell that "o1" is not a real Order. It has no operation with which to inspect the object. It can only pass the value back through the vocabulary supplied by the host application. For this test, identity is all the domain object needs. We Are Testing the Conversation A BUBAS business program contains decisions and orchestration. Algorithms, persistence, infrastructure, and domain-object implementations remain in Java. Its unit test should therefore concentrate on questions such as: Which domain operations were invoked?With what arguments?What values did those operations return?Which branch did the program select?Which operations were deliberately not invoked?What result did the program produce? In the example, we do not test how ORDER_TOTAL calculates a total. That belongs in the Java test for the implementation of ORDER_TOTAL. We test what the business program does when ORDER_TOTAL reports 1500.00. This division gives us two focused tests rather than one oversized test: Java tests verify the individual domain operations.BUNIT tests verify how a business program coordinates them. The BUNIT test documents the business scenario directly. An order identified by 42 exists, its total is 1500.00, the approval limit is 1000.00, and the program must not approve it. The test does not explain how to manufacture an object graph that produces those facts. Opacity Buys Mockability Mocking domain objects in a general-purpose language is often difficult precisely because the production code can observe so much about them. It may call methods, inspect nested objects, compare values, serialize the object or pass it to code that expects a particular implementation. A substitute must reproduce every observable property used along the tested path. An opaque BUBAS value has only the observations provided by the registered vocabulary. If the vocabulary exposes ORDER_TOTAL, then the mock controls the answer to ORDER_TOTAL. If it does not expose the customer’s internal account object, neither the program nor the test needs to know that such an object exists. The object boundary and the testing boundary are the same boundary. This is stronger than merely saying that business programs should avoid inspecting domain objects. They cannot inspect them unless the embedder deliberately provides an operation that does so. Consequently, a token can stand in for any opaque value as long as the mocks define how the exposed operations respond to it. Multiple objects require only multiple identities: Plain Text "LOAD_ORDER" WITH ARGS(42) RETURNS "o1" "LOAD_ORDER" WITH ARGS(43) RETURNS "o2" "ORDER_TOTAL" WITH ARGS("o1") RETURNS 1500.00 "ORDER_TOTAL" WITH ARGS("o2") RETURNS 200.00 The test describes the distinctions that matter without constructing either order. The Test Uses the Real Language A dangerous form of mocking creates a second, simplified interface used only by tests. Eventually, the production vocabulary changes while the test vocabulary does not. BUNIT does not compile the business program against a parallel language. The program under test is compiled against the real sealed BUBAS language. Mocking happens later, at dispatch. Therefore, the test cannot silently keep using an operation that no longer exists in the production language. Nor can it casually return a value of the wrong BUBAS type. Before executing a test, BUNIT checks the mocks and the test configuration. It can report problems such as: A mock declared with the wrong number of arguments;A mock returning a value incompatible with the real operation;An argument supplied for a parameter the program does not accept;A mocked command that should initialize a variable but does not provide its value. The test reports these errors before the business program runs. The test remains artificial — as every unit test is — but it is artificial inside the actual language contract. Do Not Assert Everything A test becomes fragile when it records every interaction, whether or not that interaction matters to the scenario. BUNIT allows an expectation to specify only the relevant part of a call. For example: Plain Text "LOG_EVENT _, _" WAS CALLED WITH ARGS("INFO", CONTAINS("over limit")) The test requires an informational log message containing "over limit". It does not require the complete message to remain byte-for-byte identical. Similarly: Plain Text "APPROVE _" WAS NOT CALLED expresses the important negative requirement without inventing an Order merely to compare it with another Order. The purpose is not to reproduce the execution trace. It is to state the observable facts that define the business case. What This Does Not Test Opaque tokens do not prove that the Java implementation of LOAD_ORDER returns the right order. They do not prove that ORDER_TOTAL calculates taxes correctly or that APPROVE commits a transaction. Those operations require their own Java unit and integration tests. BUNIT tests the program at the orchestration boundary. This makes it possible to test business decisions without databases, service containers, or complete domain-object graphs, but it does not replace testing below or beyond that boundary. Nor does BUNIT make every Java application automatically testable. The application developer first has to expose a suitably designed vocabulary. If one enormous operation performs loading, calculation, approval and notification internally, BUNIT can mock that operation but cannot test the decisions hidden inside it. Testability therefore provides feedback about vocabulary design. Operations should represent meaningful domain capabilities at the level where business programs genuinely make choices. The Deeper Result Opaque domain types may initially look like a limitation. The program cannot examine its own values freely. It has to ask the vocabulary to interpret them. That limitation creates a clean separation: Java owns domain representation and implementation.BUBAS owns orchestration and decisions.BUNIT replaces domain capabilities at that same boundary.Tokens replace complex objects with identity when identity is all the test requires. The production program becomes independent of domain-object structure. The unit test inherits that independence. We do not need a fake Order with a fake customer containing fake line items whose prices happen to add up to 1500.00. For this business decision, we need only to say: Plain Text "ORDER_TOTAL" WITH ARGS("o1") RETURNS 1500.00 The business program never needed to know what was inside the order. Neither does its test. The detailed code and the BUBAS framework are available as open source at https://github.com/verhas/bubas.
One of the most important parts of API test automation is validating the response body to ensure data integrity. This step plays a key role in functional API testing, as it helps confirm that the API is returning the right data in the expected format. Response body validation isn’t limited to a specific request type; it applies equally to POST, GET, PUT, and PATCH APIs. The same validation approach can be used for any API response to verify the data returned by the service. Playwright offers multiple ways to validate response bodies. In this tutorial, I’ll walk you through these approaches to help you efficiently perform assertions on the response data using best practices. Checkout the previous tutorial blog to learn about Installation, the demo application, and how to send GET API requests with Playwright. How to Verify the Response Structure Response structure checks ensure that an API consistently returns data in the expected format, protecting the contract between backend services and their consumers. They help catch breaking changes early, such as missing or renamed fields, even when the API still returns a successful status code. TypeScript test("GET Order details and perform structure check", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody).toHaveProperty("message"); expect(responseBody).toHaveProperty("orders"); expect(responseBody.orders[0]).toHaveProperty("id"); expect(responseBody.orders[0]).toHaveProperty("product_name"); }); This test focuses on validating the structure of the API response. It validates that the response body contains the expected top-level keys and that each order object includes the required fields. Basic Assertions The basic assertions validate API success and data presence, making them a good first layer of verification before deeper structure or data-level checks. TypeScript test("Get order details and perform basic level verification", async ({ request, }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(Array.isArray(responseBody.orders)).toBeTruthy(); expect(responseBody.orders.length).toBeGreaterThan(0); }); This test performs a basic level check to confirm that the endpoint works as expected and returns the expected data in the response. After parsing the response body, the assertions focus on the following essential basic-level checks: TypeScript expect(responseBody.message).toBe("Order found!!"); The above line of code verifies that the API returns the expected message text in the response body. TypeScript expect(Array.isArray(responseBody.orders)).toBeTruthy(); This line of code ensures that the orders field in the response is an array, validating the basic response format. TypeScript expect(responseBody.orders.length).toBeGreaterThan(0); This part of the test confirms that at least one order is returned in the orders array, ensuring the response contains required data. How to Verify Response Data With Details Validating the actual data returned in the response is essential to ensure that the API response contains the correct values. TypeScript test("Get order and verify order details", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); expect(order.id).toBeDefined(); expect(order.user_id).toEqual("1"); expect(order.product_id).toEqual("79"); expect(order.product_name).toEqual("5 star 10gm Chocobar"); }); The following code ensures that the response has a valid identifier and it is not missing or empty. TypeScript expect(order.id).not.toBeNull(); expect(order.id).toBeDefined(); This check is required because the API generates the order ID when a new order is created in the system. It ensures that the “id” field has a valid value generated and assigned to it, since this “id” is used to retrieve, update, or delete order data. TypeScript expect(order.user_id).toEqual("1"); expect(order.product_id).toEqual("79"); expect(order.product_name).toEqual("5 star 10gm Chocobar"); These statements assert that the order details are retrieved correctly for the respective request. The “user_id” - “1” was sent in the request, and verifying it in the response, along with the other order details such as “product_id” and “product_name,” ensures that the correct data is returned. How to Verify Response Data by Matching Objects and Arrays Playwright allows response data verification by matching objects and arrays partially within the API response. This approach is useful because it makes tests more flexible and confirms that the API returns the correct data structure and values. TypeScript test("Get order and verify matching object and array", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody).toMatchObject({ message: "Order found!!", orders: expect.arrayContaining([ expect.objectContaining({ product_id: "79", product_name: "5 star 10gm Chocobar", product_amount: 5, qty: 1, tax_amt: 0.5, total_amt: 5.5, }), ]), }); }); In this test, the toMatchObject assertion verifies that the response contains a “message” with the expected value “Order found!!” and an orders array. Within the array, "expect.arrayContaining" ensures that at least one order matches the expected data, while "expect.objectContaining" verifies only the values in the specified fields of that order. Using Best Practices to Perform Assertions Best practices create stable, maintainable API automation tests by combining basic checks with flexible data matching. TypeScript test("Get Order details API test with best practice", async ({ request }) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(responseBody.orders.length).toBeGreaterThan(0); expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); }); The test sends a GET request to fetch order details for “user_id”-“1". The use of failOnStatusCode: true ensures the test fails immediately if the API does not return a 2xx status code. The response is then parsed into a JSON object for validation. The assertions are structured in layers: TypeScript expect(responseBody.message).toBe("Order found!!"); This assertion verifies the message text, confirming that the API returns the correct message when an order is found. TypeScript expect(responseBody.orders.length).toBeGreaterThan(0); This statement ensures meaningful data is returned and avoids false positives when the array is empty. TypeScript expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); The final part of the code performs the final assertion using arrayContaining and objectContaining to verify that at least one order has the expected “id” and “product_name”, without asserting every field. These layered validations improve clarity by verifying structure, data presence, and key data values in sequence. Extracting Data From the Response Extracting data from the API response is a common and widely used pattern in API test automation. It is important in multiple ways, such as reusing the data in further tests for dynamic testing and end-to-end validation. TypeScript test('Get order details and extract the order id', async({request}) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { id: 1, }, failOnStatusCode: true, }); const responseBody = await response.json(); expect(responseBody.message).toBe("Order found!!"); expect(responseBody.orders.length).toBeGreaterThan(0); expect(responseBody.orders).toEqual( expect.arrayContaining([ expect.objectContaining({ id: 1, product_name: "5 star 10gm Chocobar", }), ]) ); const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); const order_id= order.id; console.log(order_id); const product_name = order.product_name console.log(product_name) }); This test sends a GET API request and performs basic validations to ensure the API response is reliable. TypeScript const order = responseBody.orders[0]; expect(order.id).not.toBeNull(); const order_id= order.id; console.log(order_id); The code above extracts the “order_id” from the order object in the response. Before accessing it, an assertion is made to verify that the value is not null. Finally, the value of the order_id is printed in the console. TypeScript const product_name = order.product_name console.log(product_name) Similarly, other values, such as product_name, can also be extracted. Attaching the Response Body to the Playwright Report The Playwright report, by default, shows the steps executed, the number of tests run, pass/fail status, and time taken to run the tests. However, it does not attach the response body to the test report. Attaching the response body to the report improves visibility and makes the test report more informative and transparent. The following code shows how to extract the required metadata and attach it to the Playwright report. TypeScript test("Get order details API and attach the response details to the report", async ({ request, }, testInfo) => { const response = await request.get("http://localhost:3004/getOrder/", { params: { user_id: "1", }, }); expect(response.status()).toBe(200); const status = response.status(); const statusText = response.statusText(); const headers = response.headers(); const body = await response.json(); const fullResponse = { status, statusText, headers, body, }; await testInfo.attach("Full API Response", { body: JSON.stringify(fullResponse, null, 2), contentType: "application/json", }); }); The testInfo is a built-in Playwright fixture and provides utilities to manage and inspect test execution, such as attaching files to reports, updating test timeouts, and identifying the currently running test. The following lines of code extract the response metadata, such as the status code, status text, headers, and response body. TypeScript const status = response.status(); const statusText = response.statusText(); const headers = response.headers(); const body = await response.json(); Next, let’s combine all response details and create a single object containing: Status codeStatus textHeadersResponse body TypeScript const fullResponse = { status, statusText, headers, body, }; Finally, let’s attach these details to the report using the testInfo.attach() method as shown below: TypeScript await testInfo.attach("Full API Response", { body: JSON.stringify(fullResponse, null, 2), contentType: "application/json", }); The testInfo.attach() adds an attachment to the Playwright report. The attach() method has 3 parameters: Name of the attachment: The first parameter is the name, “Full API Response”, that will be shown for the attachment.Body of the attachment: The second parameter is for the body of the attachment. The JSON.stringify(fullResponse, null, 2) has 3 arguments. The first argument converts the fullResponse object into a readable, pretty-formatted JSON. The second argument is the replacer, which is null. It ensures that all properties from the fullResponse object are included as they are, without modifying anything. The third argument controls pretty-printing. Here, “2” means indent nested JSON by 2 spaces.Content type: This parameter ensures that the report treats the attachment as JSON. The following screenshot is generated after the tests are run: Test Execution Running the tests in Playwright is simple and easy. We can run the following command from the terminal: Plain Text npx playwright test To generate the report, the following command can be used: Plain Text npx playwright show-report Summary Playwright provides multiple approaches, including structure checks and matching objects and arrays for verifying response data. The right strategy should be chosen based on your project’s requirements. Based on my experience, combining response structure checks with response data validation, including the matching object and array strategy, can be used as an effective approach for validating API responses. Happy testing!
Artificial Intelligence has become one of the most influential technologies in modern software development. From chatbots and recommendation systems to sentiment analysis and intelligent search, machine learning models are now expected features in many mobile applications. For several years, integrating AI into iOS applications almost always meant sending user data to cloud services. APIs such as OpenAI, Anthropic Claude, and Google Gemini allowed developers to leverage state-of-the-art language models without worrying about infrastructure or hardware limitations. While this approach is simple, it also introduces several challenges, including network latency, API costs, internet dependency, and privacy concerns. Fortunately, the landscape has changed dramatically. Today's Apple devices contain incredibly powerful hardware, including the Apple Neural Engine (ANE), powerful GPUs, and highly optimized CPUs capable of running sophisticated machine learning models directly on the device. This shift has made on-device AI more practical than ever. Instead of relying entirely on cloud services, developers can now deploy transformer models directly within their applications, enabling offline functionality, lower latency, improved privacy, and reduced operational costs. In this article, we'll explore how to integrate an ONNX-based transformer model into an iOS application using Swift. We'll load a model, execute inference with ONNX Runtime, and prepare the necessary transformer inputs for models such as DistilBERT. A Brief History of LLM Integration in Swift Projects Large Language Models were originally designed to run on powerful cloud infrastructure because of their immense computational requirements. Training these models required thousands of GPUs, and even inference demanded hardware far beyond what smartphones could provide at the time. Because of these limitations, early Swift applications integrated AI almost exclusively through cloud APIs. User prompts were transmitted to remote servers where the model generated a response before sending the results back to the application. Although this architecture worked well, it also introduced unavoidable drawbacks: Internet connectivity became mandatory.Responses depended on network latency.User data had to leave the device.API usage generated recurring operational costs. Meanwhile, Apple continued investing heavily in machine learning acceleration. The introduction of the Apple Neural Engine in 2017 marked a turning point. Every generation of iPhone, iPad, and Mac became increasingly capable of executing neural networks efficiently. At the same time, Apple expanded Core ML, Metal Performance Shaders, and hardware acceleration APIs that allowed developers to run increasingly sophisticated models locally. The open-source AI community accelerated this transition even further. Frameworks such as llama.cpp, MLX, MLC LLM, and ONNX Runtime made it possible to execute optimized transformer models directly on Apple devices. Developers could now deploy popular open-source models including Llama, Mistral, Phi, Gemma, and Qwen without requiring any cloud infrastructure. Apple later introduced Foundation Models as part of Apple Intelligence, further demonstrating the industry's movement toward local AI processing. Today, Swift developers have more choices than ever before. Depending on the application, developers can choose between cloud-hosted models, hybrid cloud/local inference, or fully offline on-device inference. For many applications, including text classification, semantic search, recommendation engines, and lightweight AI assistants, on-device inference has become the preferred solution. Why ONNX? Before diving into the implementation, it's worth understanding why ONNX has become one of the most popular deployment formats for machine learning models. ONNX (Open Neural Network Exchange) is an open standard for representing machine learning models. Instead of locking your project into a specific framework such as TensorFlow or PyTorch, ONNX provides a portable format that can be executed across many different platforms. This portability offers several advantages. A model trained in Python using PyTorch can be exported as an .onnx file and later executed inside an iOS application without rewriting the model itself. Likewise, the exact same model can often be shared between iOS, Android, Windows, Linux, and macOS. This dramatically simplifies deployment across multiple platforms. Microsoft maintains ONNX Runtime, a highly optimized inference engine capable of executing ONNX models efficiently across different hardware accelerators. For Swift developers, this means we only need to load the ONNX model, provide the expected inputs, and retrieve the outputs generated by the runtime. Loading an ONNX Model Let's assume we've already trained our transformer model and exported it to ONNX. Our project now contains a file named MoodClassifier.onnx. The model can either be added directly to the application bundle or packaged as a Swift Package resource. The first step is locating the model inside the application. Swift let modelPath = Bundle.main.path(forResource: "MoodClassifier", ofType: "onnx") If modelPath is not nil, the application has successfully located the model. If it returns nil, verify the following: The model has been added to the target.The filename matches exactly.The resource exists inside the application bundle.The file extension is correct. Successfully locating the model is the first indication that everything has been configured correctly. Installing ONNX Runtime Executing an ONNX model requires an inference engine. Fortunately, Microsoft provides an official Swift Package for ONNX Runtime that can be added using Swift Package Manager. Swift .package(url: "https://github.com/microsoft/onnxruntime-swift-package-manager",from: "1.24.2") After adding the dependency, we're ready to create an inference session. Running Inference Running inference simply means executing a trained machine learning model using new input data. Unlike training, inference does not modify the model. It only computes predictions. Creating an ONNX Runtime session requires three primary components: ORTEnvORTSessionOptionsORTSession The environment configures runtime behavior and logging. The session options allow developers to customize execution behavior. Finally, the session loads the model into memory and prepares it for inference. Swift let env = try ORTEnv(loggingLevel: .warning) let options = try ORTSessionOptions() let modelPath = Bundle.main.path(forResource: "MoodClassifier", ofType: "onnx")! let session = try ORTSession(env: env, modelPath: modelPath, sessionOptions: options) let outputs = try session.run(withInputs: [:], outputNames: ["logits"], runOptions: nil) In this example, the model returns a tensor called logits. Depending on how the model was exported, your output tensor may have a different name. Always inspect the exported model to determine the available output names. Preparing Inputs for Transformer Models Most transformer models, including DistilBERT, BERT, and RoBERTa, cannot process raw text directly. Instead, they expect numerical tensors representing the input sentence. This process is called tokenization. Tokenization converts natural language into token IDs that correspond to entries within the model's vocabulary. Alongside the token IDs, transformer models also require an attention mask. The attention mask tells the model which tokens belong to the original sentence and which tokens are merely padding added to maintain a fixed sequence length. Using the correct tokenizer is extremely important. The tokenizer used during inference must be identical to the tokenizer used while training the model. Even small differences in vocabulary or preprocessing rules can generate completely different token IDs, resulting in poor predictions despite using the correct model. Using Swift Transformers Hugging Face provides an excellent package called Swift Transformers that simplifies tokenization directly within Swift. The package can be added using Swift Package Manager. Swift .package(url: "https://github.com/huggingface/swift-transformers", from: "1.3.3" ) Once installed, you can load the tokenizer that matches the model used during training and generate the input_ids and attention_mask required by the transformer. After generating these arrays, they must be converted into ONNX tensors before inference. Creating ONNX Input Tensors The generated token arrays must be wrapped inside ORTValue tensors. The following example converts both arrays into tensors before executing the model. Swift let env = try ORTEnv(loggingLevel: .warning) let options = try ORTSessionOptions() let modelPath = Bundle.main.path(forResource: "MoodClassifier", ofType: "onnx")! let session = try ORTSession(env: env, modelPath: modelPath, sessionOptions: options) let inputData = NSMutableData(bytes: &inputIDs, length: inputIDs.count * MemoryLayout<Int64>.size) let inputTensor = try ORTValue(tensorData: inputData, elementType: .int64, shape: [1, inputIDs.count] as [NSNumber]) let attentionData = NSMutableData(bytes: &attentionMask, length: attentionMask.count * MemoryLayout<Int64>.size) let attentionTensor = try ORTValue(tensorData: attentionData, elementType: .int64, shape: [1, attentionMask.count] as [NSNumber]) let outputs = try session.run(withInputs: ["input_ids": inputTensor, "attention_mask": attentionTensor], outputNames: ["logits"], runOptions: nil) In this example, two tensors are created: input_ids, which contains the numerical representation of the input text, and attention_mask, which tells the model which tokens should participate in the attention mechanism. These tensors are then passed into the ONNX Runtime session, which executes the model and returns the requested outputs. Conclusion On-device AI is no longer a niche capability reserved for flagship applications. Thanks to frameworks such as ONNX Runtime and Swift Transformers, integrating transformer models into iOS projects has become both accessible and practical. In this article, we explored how to load an ONNX model, execute inference using Microsoft's ONNX Runtime, and prepare the input_ids and attention_mask tensors required by transformer models. These components form the foundation for deploying a wide range of AI-powered features directly within Swift applications. As Apple's hardware continues to evolve and transformer models become increasingly efficient, local inference will play an even greater role in the future of mobile development. Whether you're building a sentiment analyzer, semantic search engine, recommendation system, or lightweight AI assistant, ONNX Runtime provides a robust and portable solution for bringing modern machine learning to iOS.