Evaluating Flexday AI
Responsible AI
How Flexday AI controls its AI and what people decide: publishing, evaluation, grounding, guardrails, tool allow-lists and budgets, mapped to the NIST AI RMF.
- Functional users
- Technical
Last reviewed
Flexday AI uses AI models in two places: the Builder, which writes a Solution from a description, and Agents, the assistants those Solutions contain. In both, code sets the limits the model works within. What an Agent may do is an explicit list of grants, and what reaches its model is checked first. The Builder works only when a person asks it to in chat, but not everything it does waits for a person: it can publish the Flows and Agents it builds, and its data, schema, saved-query and gateway changes apply as it makes them, while changes to an app's files wait in a Draft until a person deploys them (the first build deploys itself). What takes effect, and when sets this out. Every completed model call is recorded, and the usage pages estimate its cost. This page describes those controls the way an AI risk reviewer would test them, and maps them to the NIST AI Risk Management Framework. The control-by-control detail is on AI safety.
At a glance
- Clear about what goes live. Agents and Flows change on a draft, and only a publish creates the version people use; the Builder can publish the Flows and Agents it builds, and its data, query and gateway changes apply as it makes them (what takes effect, and when). A conversation keeps the version it started with.
- Tested before release. An Agent can be required to pass its evaluation set before each new version is published.
- Limits are enforced in code, not asked of the model. Tool grants are an allow-list, input checks run before the model sees a message, and budgets are enforced by the engine.
- Answers you can check. An answer lists the documents its wording matches, and an Agent can be told to answer only from connected knowledge.
- Nothing hidden. The Builder narrates every action, must declare anything it did not build for real, and every conversation and every completed model call is recorded.
Human control points
| Control point | What happens |
|---|---|
| Drafts and publishing | An Agent or a Flow is edited on a draft that serves nobody. Publishing checks it, then snapshots a numbered version, and that version is what people and systems use. A person publishes, or the Builder publishes the Flows and Agents it builds, through the same checks. There is no unpublish: an Agent is switched off instead. |
| Deploying an app | The first real build of an app deploys itself so there is something to open; every later change waits in the draft until a person deploys it. |
| Pinned versions | A conversation with a published Agent keeps the version it opened with (a Playground conversation on the draft follows each edit), and a Flow run keeps the version it started with, so a publish never changes the rules part way through. |
| "Don't build yet" | When the person's words put the Builder on hold, every tool that would change something is refused in code for that turn, except saving the Builder's own notes; the Builder can still read and answer. |
| Questions back to the person | The Builder and an Agent can both stop and ask a question instead of guessing: an Agent's turn waits for the answer, and the Builder ends its turn and carries on from your reply. |
| Proposals instead of writes | Where the Builder needs a setting only a person should create, such as a Variable, it proposes it and stops there; the Variable exists only once the person clicks Done. |
Evaluation before release
An Agent's Evals hold golden test cases. Each case is a question with assertions about the answer: text it must or must not contain, a pattern, a JSON schema, tools it must or must not call, limits on tool calls and tokens, and a rubric scored by a model. A run sends every case through the real engine, in isolated test conversations.
An Agent can be set to Require evals to pass before publishing: publishing is then refused unless the default evaluation set's latest run passed against the current draft, unchanged since the run. The setting is off until you switch it on for an Agent. Evaluation conversations are left out of analytics (they are not real use) but counted in usage (they cost money).
Grounding and citations
- Citations are decided by code. After an answer, deterministic rules compare its wording with the passages the Agent retrieved and list the matching documents as sources. No model call decides what is cited.
- Answer only from connected knowledge. With this option an Agent is told to search its sources before stating a fact about your business. If it answers without consulting any, it is reminded once. The reminder is a nudge, not a block, and publishing refuses the option for an Agent that has no knowledge source granted.
Guardrails
Input checks run before the Agent's model sees a message, cheapest first: rate, length, blocked patterns, personal data (block, mask or warn), blocked topics, then optional moderation. Blocked topics described in words and moderation send the message to a separate fast model to decide. A message a guardrail blocks gets a refusal and is not passed to the Agent's model on that turn; it stays in the stored conversation as it was sent, and later turns include it in the recent history they give the model.
- Block or flag. A blocked topic either refuses the message or lets the conversation continue and records it. A topic described in words can also start a Flow, whichever its action, for example to alert HR, and a flagged topic is never announced to the person it is about.
- Fail closed where you said block. A check set to block refuses the message if the check itself cannot run. A check set only to warn or flag lets the message through.
- Masking after the model. When personal data is set to mask or block, it is masked in the Agent's reply, as shown and as stored.
Tools as an allow-list
An Agent can use only the tools it is granted, from ten kinds: Fact Base queries, document search, File Store access, running a Flow, delegating to another Agent, HTTP requests, email, channel messages, handing off to a person, and memory. A few built-in conversation tools, such as asking a question or showing a form, come from the Agent's own settings instead. Every resource a grant names must be in the same Solution, which is checked when a conversation starts. No wording in a message can add a tool.
- A Fact Base grant can name one of the Fact Base's own database roles, and a grant that would bypass the Fact Base's row policies for end users is refused at publish.
- HTTP requests must start with an allow-listed address prefix, and no hop may reach a private network.
- An Agent sends at most one email per turn, under the Solution's email rules.
- Writing files is never implied by a File Store grant; it must be granted on its own.
Budgets and ceilings
Every Agent turn has budgets for tool calls, output tokens and elapsed time. The defaults are 15 tool calls, 32,000 tokens and two minutes; an Agent may raise them only up to the platform's ceilings of 30 tool calls, 64,000 tokens and ten minutes. A conversation has a turn limit as well (200 by default, at most 1,000). The engine checks the budgets between steps: at 80% the Agent is told to wrap up, and at 100% it gets no more tools and gives a short final answer. Delegation to specialist Agents shares one allowance per turn.
Prompt-injection containment
Text an Agent reads from tools and documents is not trusted as instructions:
- Tool results and retrieved passages are wrapped and marked as untrusted, and the Agent is told to treat anything inside them as information, not orders.
- The real limit is the grant list. Injected text cannot add a tool, reach another Solution's resources, or send email beyond the per-turn cap.
- An Agent can never read a Secret Variable, and the Builder's tools never show a stored credential to the model.
Transparency
- Every Builder action is narrated. The build log has a line for every action the Builder takes, including what it reads.
- Placeholders are declared. Every finish carries the Builder's declaration of what is still a placeholder, and the Solution's overview lists it under Not built for real yet. A request for something only the platform can build, such as an Agent or a Doc Base, must be met by a real resource or named in that declaration, or the Builder cannot finish.
- "Done when" is checked. Acceptance criteria you give the Builder are checked by a separate model call against an inventory the platform reads of what the Solution contains, and the result is shown to you as met, not met or can't verify.
Model and provider choice
Models are chosen per use case from the platform's model catalog: Anthropic Claude models, and GPT models through Azure AI Foundry where a deployment is configured for them. Workspace owners and admins can override the model for each use case listed on the workspace's AI models page, and an Agent can name its own model; other AI features, such as guardrail checks, use the deployment's default. Those settings offer only Anthropic models for the Builder, because its turn-by-turn checks are built around them; Agents may use either provider. A dedicated deployment uses the model-provider keys it is configured with, which can be your own. See AI models.
Monitoring
| Record | What it shows |
|---|---|
| Agent Analytics | Real use only: conversations, turns, distinct people, tool use, sources cited, guardrail events and feedback |
| Agent Sessions | The conversations themselves, turn by turn, for review |
| Usage | Every completed model call, with tokens and an estimated cost, attributed to the Solution and resource it served |
| Flow runs | Every step's input, output, status and timing |
Alignment with the NIST AI RMF
Important
This table shows how platform mechanisms support the four functions of the NIST AI Risk Management Framework (AI RMF 1.0). It describes alignment, not certification: the framework is voluntary and describes outcomes for your organisation, and applying it remains your organisation's work.
| Function | What it asks for | Platform mechanisms that support it |
|---|---|---|
| Govern | Accountability, roles and policies for AI use | Workspace and Solution roles; drafts and explicit publishing; the audit trail of who changed what; workspace AI model settings; staff support access controlled by a workspace switch (on by default) that the owner can turn off, with owners emailed whenever a support session starts |
| Map | The context, purpose and risks of each AI system | Each Agent's instructions, intents and explicit tool grants; blocked topics; audience tags on knowledge; one Solution per purpose, with every resource it may reach inside it |
| Measure | Testing, evaluation and monitoring | Evaluation sets with assertions and a model-scored rubric; Agent analytics; usage and estimated cost per call; citations; feedback on answers |
| Manage | Responding to risks and changing course | The publish gate on evaluations; block and flag guardrails that can start a Flow; budgets; switching an Agent off; pinned versions and publishing forward; handoff to a person |
What stays your responsibility
- Deciding which tools and knowledge each Agent is granted, and which model provider it uses.
- Writing instructions, blocked topics and evaluation cases that reflect your policies.
- Telling the people who use an Agent what it is for and what it is not.
- Reviewing analytics and conversations, and acting on flagged topics.